Key takeaway
AI email marketing means a model inside the programme: drafting copy, scoring contacts, choosing send times, and in some products running the steps. We've sent more than 6 billion emails across two companies, and I've never watched a programme fail for lack of drafts. Gmail asks bulk senders to keep spam complaints under 0.30%, and a model that writes faster only reaches that ceiling sooner. The question isn't what it can write. It's how much of the programme it's allowed to run.
What AI email marketing actually means.
Two different products hide under the one phrase. One produces content: subject lines, body copy, images, and forty versions of a call to action. The other changes what the programme does: who gets the message, when it leaves, which version wins, and whether the send happens at all. Most guides define the term by listing both and then never separate them again. Almost everything shipping under that name today is a button inside a dashboard. The model writes, a person still does every operation around it, and what gets saved is the typing rather than the work.
I've spent over a decade building email platforms, and the pattern hasn't shifted: the teams who get something out of a model are the ones who already knew which decision they were handing it. None of the sending I've overseen in that decade ever moved on better adjectives.
The question the definition skips is who is still clicking afterwards. A tool can hold every capability on the list and leave your operating cost exactly where it was, because a suggestion somebody has to read, judge, and paste is still a job for a person.
Outbound campaigns often turn into a volume game that damages a company reputation. Startups want a better way to reach prospects without looking like spammers. The market demand is clear on The Breakout CEO because "these people are trying to build a great brand as well as their startup."
The five jobs a model does in an email programme.
Strip the marketing off and five jobs are left.
- Drafting: body copy, subject lines, and as many variants of each as you'll agree to read.
- Variant testing: generating the alternatives and calling a winner once enough of the send has landed.
- Send-time selection: choosing a per-contact hour out of when that person has opened before.
- Scoring and segmentation: ranking contacts on likelihood to buy, churn, or engage, and proposing a group from the ranking.
- List hygiene: flagging addresses that stopped engaging before they cost you a complaint.
What matters more than the count is the property all five share. Every one is a suggestion computed against your own data, so the same feature does real work on one list and nothing on another. Send-time selection needs months of opens recorded against each contact. Run it on a list you imported last month and there's no history to read, so it returns an hour that means nothing, and the screen looks the same either way. Four of the five need history you don't have on day one, which is why I'd start any of this by checking what's on the contact record rather than by comparing feature lists.
Personalization is a data problem before it's a model problem.
Personalization at scale means the content changes per contact against something you genuinely know about them. The model is the cheap half of that, because it's the half you can buy on a Tuesday.
A model with no event history personalizes on a first name, which is the 2009 version of this and the main reason teams conclude that AI did nothing for them. The precondition is unglamorous and specific: events that record what a contact did and when, purchase or usage history, custom fields that are filled rather than merely defined, and a suppression state that's current. Give a model those four and it can write a paragraph that only makes sense to one person. Give it a name and a signup date and it writes the email you already had, faster.
So the first month is instrumentation, not evaluation. Pick the two or three events that describe your product being used, get them landing on the contact record, and let the model have nothing to say for a week. None of that work is a model problem and nobody sells it as a product, which is exactly why it keeps getting skipped in favour of a tool.
Machine learning models require clean inputs to produce useful outputs. A model cannot generate relevant product recommendations without tracking user behavior across your store. We feed our system a specific diet on SaaS District that "uses customer data, product data and the on-page stuff" to constantly improve the revenue per email.
Drafting, deciding, and executing are three different jobs.
What an AI feature can produce says less about it than how far past your keyboard it's allowed to act, and that reach comes in three levels.
- AI that drafts: it writes the copy and a person publishes it, so a human reads every output before a recipient does.
- AI that decides: it picks the send time, the segment, or the winning variant, and a person approves the campaign built around that choice.
- AI that executes: it creates the campaign, builds the segment, and sends, with nobody in the path.
Nearly every product sold as AI email marketing today does the first and is described in the language of the third. The distinction is worth holding onto, because each level wants a different guardrail. Drafting wants an editor. Deciding wants a measurement window long enough to tell whether the decision was any good, since a bad send-time model raises no alert and shows up as a slightly worse quarter. Executing wants limits, an approval step, and a blast radius you chose in advance rather than discovered.
The level isn't a product tier and you don't pick one for the whole programme. It's a choice per workflow, and in my experience most teams should be happy letting a model run the welcome email and nowhere near the Black Friday send.
Automation makes it tempting to blast every contact in your database. This approach ignores user engagement and treats every subscriber as identical. The penalty is severe on Designwave when you start "sending to full lists instead of segments" and eventually destroy your sender reputation.
What AI-generated volume costs you at the mailbox provider.
A model makes sending cheaper. Almost everything that used to make sending expensive was the part that protected your reputation: the hours that forced somebody to ask whether this send was worth doing, to this many people, this week.
The receiving side sets the floor, and it doesn't care how the mail was written. Gmail treats anyone sending more than 5,000 messages a day to Gmail addresses as a bulk sender, requires SPF, DKIM, and DMARC on the sending domain, expects marketing mail to carry a one-click unsubscribe, and asks for a spam complaint rate under 0.30%, with 0.10% as the number to actually run at, in Google's sender guidelines. The unsubscribe half is a published standard rather than a vendor preference, defined in RFC 8058: a header pair, a POST from the mailbox provider, and no confirmation page in between.
A programme that triples its output against a list that didn't grow moves exactly one of those numbers, and it's the complaint rate. Three times the mail to the same people is three times the chances that somebody reaches for the spam button rather than the unsubscribe link, and a complaint rate is a ratio, so the extra volume in the denominator doesn't rescue you.
So read the complaint rate weekly once a model is writing. It's the one number here with a published ceiling, the drift is visible in Postmaster Tools well before anything gets blocked, and once placement moves you're repairing a reputation rather than fixing a send. At my last company, I had it auto-flag any list where unsubscribe and complaint rates crossed 0.5%: industry averages for both sit closer to 0.02%, so 0.5% was already a fire, not a warning.
Holiday sales events bring out the worst instincts in retail marketers. Merchants frequently decide to resurrect inactive contacts just to maximize their sending volume. The math is brutal on In the Ring with SUMO Heavy when you blast a "dusty old list of a hundred thousand people who gave me their email about four years ago" because you will immediately ruin your deliverability.
The subject lines a model will write that you shouldn't send.
A model optimising for opens will write the misleading line, because the misleading line performs. Ask for forty subject lines ranked on predicted open rate and something shaped like "Re: your order" turns up near the top of a promotional send about a sale. Nothing in the ranking knows the email isn't a reply, and the people who open it find that out in a second.
That's a complaint-rate problem before it's anything else. A subject line that misdescribes the email buys the open and spends the trust, and the spam button sits next to the reply button. Against a 0.30% ceiling and a 0.10% working number, one misdescribed send to a large segment is enough to move a ratio that small.
Three checks, and they run in this order:
- Does the subject line describe what's actually in the email?
- Does every claim in the body survive being checked against something true?
- Does the unsubscribe still work in the template the model rewrote?
Each one takes a minute, and almost nobody runs them, because a review habit built for four drafts a week doesn't survive a tool that produces forty in an afternoon. The fix is to generate fewer variants, not to read faster.
Where the model sits, and where to start.
On every mainstream platform the AI is a feature inside the interface. You open the tool, find the screen, press the button, read the suggestion, and do the rest by hand. That's the shape of the whole category, and it's why the demos are interchangeable. Every one of those platforms was designed for humans clicking buttons, and retrofitting an agent onto a dashboard-first product is bolting a motor onto a bicycle.
Most vendor models write decent copy, so the test that matters isn't the writing. Ask plainly whether something outside the interface can create the campaign, build the segment, confirm the sending domain is authenticated, and send. On most products some of that is an API and some of it is a screen, and the screen is reliably the step that hands the job back to you. An agent that can operate the platform can also send the wrong thing to everybody, so the argument that asks for execution asks for limits and an approval step in the same breath.
Four steps, in order, and the first is the one that usually gets skipped:
- Pick one workflow and one level of autonomy for it rather than switching AI on across the programme.
- Check that the data that workflow needs is already on the contact record.
- Run the model against a segment small enough to be wrong in.
- Measure it against the send it replaced, on the same segment and the same window, rather than against somebody else's benchmark.
Opens are the wrong scoreboard once a model writes the subject lines, since that's the number it optimises and it can move without selling anything. Replies, clicks to a page that matters, and revenue per thousand sent are harder to fake.
We built Nitrosend the other way round: MCP-first, so campaigns, flows, contacts, templates, and segments are an API endpoint and a CLI command before they're ever a screen, with BYO sending keys on Pro and above and unlimited contacts on every plan including Free. If your model can write the email but can't run the programme, that's the gap it doesn't close. Start free, point one message type at Nitrosend, and judge it on a week of real numbers against the send it replaced.
Sources
- Google, Email sender guidelines: the 5,000-a-day bulk-sender threshold, the SPF, DKIM, and DMARC requirement, one-click unsubscribe on marketing mail, and the 0.30% and 0.10% spam-rate numbers.
- RFC 8058: Signaling One-Click Functionality for List Email Headers, the header pair and POST behind the one-click unsubscribe requirement.
Common questions
It's the point where email software stops following a rule somebody typed and starts making the call itself. A rules engine sends at 9am because you set 9am. A model picks the hour out of when that contact has opened before, ranks who on the list is worth mailing, and writes the message. On most platforms it stops at the suggestion and hands the campaign back to a person to run, which is the difference between software that advises and software that operates the programme.
It can produce one. Whether it should send one is a different question, and the answer changes per workflow: a welcome email is a reasonable thing to hand over, and a promotional send to the whole list isn't. A model that drafts puts its output in front of a person before a recipient sees it. A model that sends puts it in front of your list first, so that's the one that needs a limit and an approval step.
It replaces the typing, not the decision about who gets the message and why. A model can generate forty subject lines in a second and has no view on whether the campaign should go out at all, which segment deserves it, or what a rising complaint rate means for next month's sending.
Four inputs, and the model isn't one of them: a behavioural event stream showing what each contact did and when, a record of what they've bought or used, custom fields that actually carry values rather than sitting defined and empty, and a suppression state that reflects today rather than last quarter. Without those, a model personalizes on a first name, which is the version of personalization everybody already had.
Not by being AI-written. It hurts deliverability by making sending cheap, so a programme sends more to a list that didn't grow and the spam complaint rate climbs. Gmail asks bulk senders to keep that rate under 0.30%, and under 0.10% is the number worth actually running at.
There's no general duty to label it, and no mailbox provider treats model-written copy differently. What still applies is what always applied: the subject line has to describe what's in the email, the claims in the body have to be true, and the unsubscribe has to work in whatever template the model rewrote. A misdescribed subject line buys the open and pays for it in complaints, and Gmail's ceiling for those is 0.30%.
Against the send it replaced, on the same segment and over the same window, rather than against a benchmark taken from somebody else's list. Measure replies, clicks to a page that matters, or revenue per thousand sent rather than opens, since opens are the number a subject-line model is optimising.
Yes. Nitrosend is MCP-first, so campaigns, flows, contacts, templates, and segments are API endpoints and MCP tools before they're screens, with a REST API and a CLI beside them. An agent can create the campaign, build the segment, and send it without anyone opening the interface.