Operations skill
Cold email campaign metrics reader
Share campaign or A/B numbers and the assistant checks the measurement first, then gives a verdict with a confidence level: winner, no difference yet, or invalid comparison.
What it does
What the skill checks
- Uses emails sent as the denominator, not leads uploaded
- Reads step-1 reply rate on a fixed list instead of a blended daily rate
- Separates human replies from automatic replies
- Counts opportunities per 1,000 leads contacted
- Checks reply maturity, sample size and that arms differ in one thing
- Confirms the sent body matches the variant before crediting copy
Use it when
- Reading an A/B test between two copies or subject lines
- Reply rate dropped and you need to know if it is real
- Reporting campaign results to a client or a manager
Example
What you paste and what you get back
A: 1,200 sent, 14 replies, 2 interested. B: 1,150 sent, 9 replies, 3 interested. Which copy wins?Reply rate: A 1.17%, B 0.78%.
Opportunities per 1,000 contacted: A 1.7, B 2.6.
Verdict: no difference yet. Confidence: low.
Why: 2 versus 3 opportunities is noise, and B has fewer replies but more interest.
Next: 3,000 more sends per arm, same list and start time.Install
Install it in your assistant
The same file works everywhere. Claude and Claude Code load it as a skill; the other assistants follow it as instructions.
Claude Code
- Run the command below. It saves the skill to your personal skills folder, available in every project. For one project only, use .claude/skills in the project instead.
- Start a new Claude Code session. Claude uses the skill when your request matches its description, or when you name it.
mkdir -p ~/.claude/skills/campaign-metrics-reader
curl -fsSL https://outreach2day.com/skills/campaign-metrics-reader/SKILL.md -o ~/.claude/skills/campaign-metrics-reader/SKILL.mdClaude (web and desktop)
- Download the ZIP file.
- In Claude, open Settings, find Skills under Capabilities and upload the ZIP. Skills need code execution to be turned on for your account.
- Ask for the task in any chat. Claude loads the skill when the request matches.
ChatGPT
- Download SKILL.md.
- Paste its text into a Project's instructions or a custom GPT's instructions. For a single chat, attach the file and write: follow the instructions in this file.
Gemini
- Download SKILL.md.
- Create a Gem and paste the file's text into its instructions, or attach the file to a chat and ask Gemini to follow it.
Grok
- Download SKILL.md.
- Paste its text into a Project's instructions, or attach the file to a chat and ask Grok to follow it.
Cursor and other coding agents
- Save the file as a project rule with the command below. Cursor reads the description to decide when to apply it.
- Agents that read AGENTS.md (Codex and others): paste the text into AGENTS.md or reference the file from it.
mkdir -p .cursor/rules
curl -fsSL https://outreach2day.com/skills/campaign-metrics-reader/SKILL.md -o .cursor/rules/campaign-metrics-reader.mdcSource
The full SKILL.md
Read it before you install it. Change the rules to match your own setup.
---
name: campaign-metrics-reader
description: Reads cold email campaign results without the usual mistakes - uses emails sent as the denominator, reads step-1 reply rate on a fixed list instead of a blended daily rate, separates human and automatic replies, counts opportunities per 1,000 leads contacted, waits for replies to arrive before comparing, checks that samples are large enough, compares arms that differ in one thing, and confirms the actual sent body before crediting a copy. Returns a readout with a verdict and a confidence level. Use when the user shares campaign stats, an A/B test, a sequencer export or a dashboard and asks what is working, which variant won or why results dropped.
---
# Campaign metrics reader
You turn cold email numbers into a verdict the user can act on, and you say how sure you are. Most wrong conclusions in cold email come from the wrong denominator, a mixed blend or too few sends; you check those first.
## Step 1. Collect
Ask for, per campaign or variant:
1. Emails sent (not leads uploaded), per step if possible.
2. Replies, split into human and automatic (out-of-office, auto-acknowledgements) if the tool separates them.
3. Opportunities: replies classified as interested, meeting request or referral.
4. Bounces.
5. Start date and time of each arm, list source, mailboxes used.
6. The actual subject and body that were sent (from the sent log, not the draft).
## Step 2. Check the measurement
Report each as `OK` or `PROBLEM`:
1. **Denominator**: rates use emails sent. Leads uploaded but not yet contacted inflate the base and deflate every rate.
2. **Step mix**: a daily blended reply rate mixes first emails and follow-ups, and new lists with old ones. When the mix shifts, the blend moves with no real change. Read step 1 on a fixed list family.
3. **Human vs automatic replies**: they are separate counts. Automatic replies say the email arrived; they are not interest.
4. **Maturity**: most replies to a send arrive within a few hours. Do not read an arm in its first hours, and compare only arms that started at the same time.
5. **Sample size**: with a few hundred sends, a difference of one or two replies is noise. Ask for a few thousand sends per arm before a verdict on reply rates; opportunities need more.
6. **One difference**: the arms differ in one thing (copy, or list, or mailboxes). If two things changed, the result cannot be credited to either.
7. **What was sent**: the sent body matches the variant name. Check at least 3 sent emails per arm for rendered variables and spintax.
## Step 3. Compute
- Reply rate = human replies / emails sent, per step.
- Opportunities per 1,000 leads contacted = opportunities / unique leads contacted x 1,000.
- Bounce rate = bounces / emails sent.
- For two arms, give the difference and a rough confidence: with small counts, say "not distinguishable".
Reply rate alone is vanity: an arm with more replies and fewer opportunities is worse.
## Step 4. Verdict
One of: `winner`, `no difference yet`, `invalid comparison`, `investigate`. Add what would change the verdict (for example "2,000 more sends per arm").
## Output format
1. Measurement checks with OK or PROBLEM.
2. A table per arm: sent, human replies, reply rate, opportunities, opportunities per 1,000, bounce rate.
3. Verdict, confidence (low, medium, high) and the reason in two sentences.
4. Next step.
## Rules
- Never declare a winner on a one or two event difference.
- If placement may have changed (bounces up, automatic replies down on all arms at once), say the drop may be delivery, not copy, and suggest a placement test.
- Do not mix campaigns from different lists into one "copy" comparison.
## Example
Input: "A: 1,200 sent, 14 replies, 2 interested. B: 1,150 sent, 9 replies, 3 interested. Which copy wins?"
Output (abridged): Reply rate A 1.17%, B 0.78%. Opportunities per 1,000: A 1.7, B 2.6. Verdict: no difference yet, low confidence: 2 versus 3 opportunities is noise, and B has fewer replies but more interest. Next: 3,000 more sends per arm, same start time and list.
FAQ
Questions
What is a good reply rate for cold email?
It depends on list and offer too much for one number. Compare arms on the same list at the same time, and judge them by opportunities per 1,000 leads, not by reply rate alone.
How many emails do I need to send for an A/B test?
A few thousand per arm before trusting a reply-rate difference, more for opportunities. With a few hundred sends, one or two replies of difference is noise.
Why did my daily reply rate drop?
Often the mix changed: more follow-ups or a weaker new list in the blend. Read step 1 on one list family. If every arm drops at once and bounces rise, check delivery before the copy.
Should out-of-office replies count as replies?
No. Count human and automatic replies separately. Automatic replies show the email arrived, not that anyone is interested.
Mailboxes, warm-up and sending in one place
$2.50 a mailbox a month. DNS records are set for you, and every mailbox shows its warm-up numbers from day one.