---
name: seed-test-planner
description: Designs inbox placement (seed) test plans for cold email so the results can answer a real question. Builds a copy by domain matrix with a neutral control in every batch, respects per-domain test limits, ramps the number of tests gradually, sets polling and a results ledger, and states what difference the sample can detect. Use when the user wants to run spam or placement tests, compare copies, domains or mailboxes, or asks how many tests they need.
---

# Seed test planner

You plan placement tests that produce a decision, not a pile of scores. Every plan answers one question, has a control, and is sized so the answer is not noise.

## Step 1. Write the question

Pick exactly one per plan:

- Is this mailbox, domain or IP range healthy? (infrastructure question)
- Does copy A place better than copy B? (copy question)
- Did a change (new DNS, rest period, new copy) fix placement? (before and after question)

If the user has several, make several small plans and run the infrastructure one first. A copy comparison on broken infrastructure measures nothing.

## Step 2. Design the matrix

- **Neutral control in every batch.** A plain personal-style email (no offer, no links, a few sentences) sent from each mailbox under test. It separates the sender from the copy.
- **Copy question:** at least 3 copies by at least 3 sending domains, each copy on each domain. Fewer and copy is confounded with domain.
- **Infrastructure question:** control plus the production copy, from each mailbox or domain, on at least 2 different days.
- **Send exactly what you send in production:** same subject style, body, signature, links (or none), tracking settings, and the same sending tool or server path. A test sent a different way tests something else.

## Step 3. Respect limits

- At most about 2 tests per sending domain per day. Many tests from one domain in a day distort its results and can hurt the domain.
- Ramp the number of tests: start with about 25, check that results come back and nothing is blocked, then 75, then 150. Test tools rate-limit bursts; mass-creating hundreds at once can get tests or your IP blocked.
- Poll results once, a few minutes after sending, then again later if needed. Do not hammer the results endpoint.
- Use fresh addresses from the test provider for every test when it rotates seeds.

## Step 4. Size the sample

- A single test with 8 to 10 seeds per provider has a wide margin: roughly plus or minus 30 points. Use it for large effects only (inbox vs junk).
- To call a difference between two copies of around 20 points, plan on at least 6 to 10 tests per copy, spread over domains and days.
- Report results per recipient provider (Google Workspace, Gmail, Microsoft 365, Outlook, Yahoo); plan enough seeds on the provider that matches the user's leads.

## Step 5. Keep a ledger

One row per test, written at send time, never edited later:

| test_id | date_time | mailbox | domain | copy_id | control? | provider | seeds | inbox | spam | missing |

## Output format

1. The question in one sentence.
2. The matrix: rows = domains or mailboxes, columns = copies (control first), cells = number of tests and days.
3. The schedule: tests per day, ramp steps, polling times.
4. The decision rule written before the tests run, for example "copy B replaces A if its Microsoft inbox rate beats A by 20+ points across 3 domains while the control stays above 80%".
5. The ledger template.

## Rules

- No plan without a control.
- No copy ranking before infrastructure is confirmed healthy by the control.
- Write the decision rule before seeing results.

## Example

Input: "We have 2 new copies and 3 domains. Which copy is better on Outlook?"

Output (abridged): Question: copy A vs copy B on Microsoft 365 and Outlook. Matrix: 3 domains by {control, A, B}, 2 days, 1 test per cell per day = 18 tests, 2 per domain per day max, so run 6 per day over 3 days. Decision: pick the copy that beats the other by 20+ points on Microsoft on at least 2 of 3 domains while the control places 80%+.
