Blog · AI
— AI··11 min read

How to Use AI for Email Subject Line Testing at Scale

Joona Heinonen· Choco Media · Rovaniemi

Most email marketing teams know they should test subject lines. In practice, the testing is the first thing to drop when things get busy — because writing ten variants for every email feels like a lot when you’re already stretched. AI email subject line testing changes that equation significantly. At Choco Media, we’ve built a workflow that generates subject line variants, surfaces the ones most likely to perform, and structures our A/B tests efficiently — without the manual overhead that used to make this kind of rigour impractical.

This post is for marketing teams who send regular email campaigns and want to move from guessing to systematic testing without hiring a dedicated email strategist. We’ll cover the workflow we use, the prompt structures that produce useful variants, the signals we use to pre-filter before testing, and how to set up tests that actually teach you something.

By the end, you’ll have a repeatable process you can run on every campaign, a mental model for evaluating subject line quality, and enough practical examples to adapt it to your own list and voice.

Why Most Subject Line Testing Fails Before It Starts

The single biggest reason subject line tests don’t yield useful data isn’t low sample size — it’s that teams test variants that are too similar to each other. If your two test variants are “Our new feature is here” and “Introducing our new feature,” you’re not learning anything useful. The variables are so close that any difference in open rate is effectively noise.

Useful testing requires meaningful variation across distinct dimensions: emotional register, specificity, length, curiosity versus clarity, personalisation, question versus statement. Most teams don’t generate enough genuinely different variants to test those dimensions. They write two versions that felt different in the moment, but both sit in the same emotional neighbourhood.

AI is particularly good at generating variants across all five dimensions quickly. The issue isn’t creativity — it’s that you need a structured brief to get variants that are genuinely different, not just paraphrases.

The Brief Structure That Produces Useful Variants

The prompt matters more than the model. In client work we’ve found that a loose prompt like “give me ten subject lines for this email” produces ten versions of the same idea. A structured brief produces ten genuinely distinct options you can evaluate against each other.

Here’s the brief structure we use:

  1. Email topic in one sentence: what the email is actually about, stripped of marketing language
  2. Primary reader motivation: what does this person want to know or do? Not what you want them to do — what they want.
  3. Audience specifics: role, maturity level, size of company if relevant
  4. Tone range: where on the axis from direct/dry to warm/personal does your brand sit
  5. Length constraint: mobile preview typically cuts off around 40 characters, but some audiences on desktop have different reading habits
  6. Off-limits words or phrases: the clichés, the banned terms, the ones you’ve tested and know underperform

With that brief, the prompt becomes: “Generate 12 subject lines for this email across the following five emotional registers: curiosity, direct benefit, specificity, direct address, and narrative hook. Two or three variants per register. Don’t repeat the same opening word more than twice across the full set.”

What to Add for Better Output

Paste in three to five recent subject lines that performed well for your list, and ask the model to match the register — not the wording. This anchors the output in your voice without producing copies. If you have a specific offer or deadline, include it explicitly: AI will invent specificity if you don’t provide it, and invented specificity fails on delivery.

Pre-Filtering Before You Test: The Signals That Predict Performance

Not everything AI generates is worth testing. The point of AI is to expand the option set fast, not to hand every output to a test. Before running any variant, we run a quick pre-filter against four signals.

“The best email subject lines are specific, honest, and make one implicit promise. If you can’t identify the promise your subject line is making, readers can’t either.” — A principle we return to in client work.

The Four Pre-Filter Checks

After pre-filtering, you typically go from twelve variants to five or six worth testing. From there, you select two or three for an actual A/B test — or more if your list is large enough to run a multivariate.

Setting Up Tests That Actually Teach You Something

A test with insufficient volume is worse than no test at all, because it creates false confidence in a result that might have been noise. Our AI content creation work consistently shows that teams make significant decisions based on under-powered tests that reverse on the next send.

For subject line A/B testing, a general rule of thumb:

What to Measure Beyond Open Rate

Open rate tells you whether the subject line got attention. It doesn’t tell you whether it got the right attention. Track:

Building a Subject Line Learning Archive

The most underused part of any email program is the historical performance data sitting in your ESP. Every test you’ve run is a data point. The problem is it’s usually in a spreadsheet no one looks at, or buried in a platform dashboard that makes comparison hard.

We recommend maintaining a simple subject line archive — a running log with the subject line text, send date, list size, open rate, CTR, and a tag for the emotional register (curiosity, benefit-led, specificity, etc). After six months, you’ll have enough data to see real patterns: which registers your list responds to, whether personalisation moves the needle for your audience, whether length matters.

This archive also becomes a training input. When you brief AI for new variants, you can include the top ten performers from your archive and ask it to generate new lines that match the pattern. You’re building a preference model specific to your list, without needing any machine learning infrastructure. Our AI automation work helps teams set this kind of archive up in a way that stays usable — the graveyard spreadsheet is a common failure mode.

Tagging for Reusability

Over time, the archive surfaces your audience’s actual preferences rather than best-practice assumptions. Best practices are useful starting points. Your list is the only data that matters.

Using AI to Pre-Score Variants Before Testing

Beyond generation, AI is useful as a first-pass scoring layer. You can prompt a model to evaluate a set of variants against a rubric — specificity, promise clarity, likely emotional response — and produce a ranked shortlist before you spend testing budget.

The prompt we use: “Evaluate these 12 subject lines against the following criteria: specificity (does it make a concrete claim?), promise clarity (is it obvious what the email contains?), and register fit (does it match the tone profile below?). Score each 1–5 on each criterion. Do not factor in novelty — we want quality over cleverness.”

This produces a preliminary ranking in seconds. We don’t treat it as ground truth — AI scoring is not a proxy for real reader behaviour. But it quickly surfaces the variants worth elevating to a test versus the ones that fail on basic quality markers.

Prompt the Model to Explain Its Scoring

Ask for one sentence of reasoning per score. This surfaces quality signals you can apply manually to future variants, even without AI. The explanation often reveals something about the line that wasn’t obvious — for instance, that the implied promise is ambiguous, or that the opening word creates a register clash with the rest of the line. Used alongside SEO and content work, subject line intelligence consistently improves downstream engagement metrics.

Integrating AI Subject Line Testing into Your Regular Workflow

The goal isn’t to add a new tool — it’s to replace a slow, manual step with a faster, more structured one. Here’s how a realistic workflow looks:

  1. Draft the email body first. Subject line variants should reflect the email’s actual content, not be written before you know what you’re saying.
  2. Run the brief. Use the six-field structure above to constrain the prompt. Include two or three top performers from your archive as voice anchors.
  3. Generate twelve variants. Aim for coverage across registers, not just different phrasings of the same concept.
  4. Pre-filter against the four checks. Eliminate anything that fails on spam, voice fit, or promise-delivery mismatch. You should be down to five or six.
  5. Score the shortlist. Either manually against your rubric, or with a scoring prompt. Pick two or three for actual testing.
  6. Log the results. Add the variants, performance data, and register tags to your archive.

Total additional time per campaign over just writing two variants manually: fifteen to twenty minutes. The upside is you’re testing genuinely different hypotheses rather than minor rewrites, which means your archive accumulates real signal rather than noise.

Common Mistakes When Using AI for Subject Line Work

We’ve made most of these ourselves, and we see them consistently in client email programs:

A Practical Example: Briefing for a Product Update Email

Email topic: announcing a new reporting dashboard feature for existing customers.

Brief fields filled:

Variants AI produced across registers:

After pre-filtering: all five passed. We ran a three-way test on the specificity, direct address, and narrative variants — the ones furthest apart from each other. Specificity won on CTR. Direct address won on open rate. Narrative had the lowest unsubscribe rate. All three taught us something distinct.

When AI Subject Line Testing Doesn’t Help

This workflow is most valuable for teams sending regularly to a stable list. It’s less useful for one-off sends to tiny lists (under 2,000) where volume limitations make testing unreliable regardless of variant quality. It’s also less useful for highly transactional emails where the subject line is effectively fixed by the action that triggered the email — a password reset or order confirmation doesn’t benefit from register variation.

Where it works best: newsletters, product updates, promotional campaigns, re-engagement sequences, and onboarding emails. These have enough cadence, volume, and content variation to make systematic testing accumulate real learning over time.

If you’re building or refining your email program and want structured testing in place alongside your content work, we’re easy to reach. We work with teams who want a tighter feedback loop between what they’re sending and what their list is actually responding to.

— Work with Choco Media

Want posts like this working for your business?

10–40 SEO + AI-optimised blog posts a month, researched, senior-edited and published straight to your site. Built to rank on Google and get cited by ChatGPT, Claude and Gemini.

See plans — from €199/mo →
No start-up fee · Price locked for 12 months · Cancel any time after
← All storiesNext story →
— Free tips, monthly

Get the playbook, for free.

One short letter a month — the prompts we use, the campaigns that worked, the AI tools worth the time. No sales pitch, just field notes.

— Want us to do it for you?

Hire the agency.

AI-accelerated content, paid media, brand and web — delivered by one small team that talks to itself. Currently taking on a handful of clients each quarter.

Book a call