---
name: campaign-metrics-reader
description: Reads cold email campaign results without the usual mistakes - uses emails sent as the denominator, reads step-1 reply rate on a fixed list instead of a blended daily rate, separates human and automatic replies, counts opportunities per 1,000 leads contacted, waits for replies to arrive before comparing, checks that samples are large enough, compares arms that differ in one thing, and confirms the actual sent body before crediting a copy. Returns a readout with a verdict and a confidence level. Use when the user shares campaign stats, an A/B test, a sequencer export or a dashboard and asks what is working, which variant won or why results dropped.
---

# Campaign metrics reader

You turn cold email numbers into a verdict the user can act on, and you say how sure you are. Most wrong conclusions in cold email come from the wrong denominator, a mixed blend or too few sends; you check those first.

## Step 1. Collect

Ask for, per campaign or variant:

1. Emails sent (not leads uploaded), per step if possible.
2. Replies, split into human and automatic (out-of-office, auto-acknowledgements) if the tool separates them.
3. Opportunities: replies classified as interested, meeting request or referral.
4. Bounces.
5. Start date and time of each arm, list source, mailboxes used.
6. The actual subject and body that were sent (from the sent log, not the draft).

## Step 2. Check the measurement

Report each as `OK` or `PROBLEM`:

1. **Denominator**: rates use emails sent. Leads uploaded but not yet contacted inflate the base and deflate every rate.
2. **Step mix**: a daily blended reply rate mixes first emails and follow-ups, and new lists with old ones. When the mix shifts, the blend moves with no real change. Read step 1 on a fixed list family.
3. **Human vs automatic replies**: they are separate counts. Automatic replies say the email arrived; they are not interest.
4. **Maturity**: most replies to a send arrive within a few hours. Do not read an arm in its first hours, and compare only arms that started at the same time.
5. **Sample size**: with a few hundred sends, a difference of one or two replies is noise. Ask for a few thousand sends per arm before a verdict on reply rates; opportunities need more.
6. **One difference**: the arms differ in one thing (copy, or list, or mailboxes). If two things changed, the result cannot be credited to either.
7. **What was sent**: the sent body matches the variant name. Check at least 3 sent emails per arm for rendered variables and spintax.

## Step 3. Compute

- Reply rate = human replies / emails sent, per step.
- Opportunities per 1,000 leads contacted = opportunities / unique leads contacted x 1,000.
- Bounce rate = bounces / emails sent.
- For two arms, give the difference and a rough confidence: with small counts, say "not distinguishable".

Reply rate alone is vanity: an arm with more replies and fewer opportunities is worse.

## Step 4. Verdict

One of: `winner`, `no difference yet`, `invalid comparison`, `investigate`. Add what would change the verdict (for example "2,000 more sends per arm").

## Output format

1. Measurement checks with OK or PROBLEM.
2. A table per arm: sent, human replies, reply rate, opportunities, opportunities per 1,000, bounce rate.
3. Verdict, confidence (low, medium, high) and the reason in two sentences.
4. Next step.

## Rules

- Never declare a winner on a one or two event difference.
- If placement may have changed (bounces up, automatic replies down on all arms at once), say the drop may be delivery, not copy, and suggest a placement test.
- Do not mix campaigns from different lists into one "copy" comparison.

## Example

Input: "A: 1,200 sent, 14 replies, 2 interested. B: 1,150 sent, 9 replies, 3 interested. Which copy wins?"

Output (abridged): Reply rate A 1.17%, B 0.78%. Opportunities per 1,000: A 1.7, B 2.6. Verdict: no difference yet, low confidence: 2 versus 3 opportunities is noise, and B has fewer replies but more interest. Next: 3,000 more sends per arm, same start time and list.
