Cold Outreach A/B Test Log & Decision Tracker
A structured template for logging every cold email and LinkedIn message A/B test your team runs, complete with copy, sample sizes, results, confidence indicators, and a mandatory go/no-go verdict. Builds a proprietary, compounding record of what works for your specific ICP over time.
How to use it
Open a new spreadsheet and recreate the tables below as tabs, or paste them directly into your existing sales ops tracker. Log every test before you send it, fill in results once you hit minimum sample size, and complete the Decision field immediately so findings don't get lost. Review the Monthly Summary rows in your pipeline meeting each month.
What's inside
- Pre-test setup fields covering hypothesis, element tested, channel, and ICP segment
- Side-by-side control vs. variant copy log with character counts
- Minimum sample size guide by test type (subject line, opening line, CTA, PS line)
- Results table with open rate, reply rate, positive reply rate, and meeting booked rate per arm
- Statistical confidence indicator with a simple calculation guide and traffic-light key
- Mandatory Decision field with three options: Adopt, Discard, or Retest, each requiring written notes
- Common failure modes column with named causes such as send-day confound and list quality drop
- Monthly Summary row for manager reporting in pipeline reviews
- When NOT to run an A/B test guidance note
Purpose
Every team has opinions about subject lines. Few have evidence. This tracker turns your outreach experiments into a proprietary record, so the knowledge stays in the team when reps leave, compounds over time, and gives your manager something real to show in pipeline reviews.
This is not a grading tool for a single email. It is not a benchmark comparison. It is a running log of controlled experiments, each with a clear verdict.
When NOT to run an A/B test
Before you open a new row, check these conditions. If any apply, fix them first.
- Your send volume is under 50 per week per arm. You will not reach minimum sample in a reasonable timeframe and the result will be noise.
- Your list quality has changed between sends (different source, different scrape date, different ICP tier). The test is confounded before it starts.
- You changed more than one element between control and variant. You will not know what caused the difference.
- You are mid-sequence. Test on step 1 first. Multi-touch sequences amplify noise from early steps.
Minimum Sample Size Guide
Use this before you start a test, not after. These are floors, not targets. Aim higher if you can.
| Element Tested | Minimum per Arm | Rationale |
|---|---|---|
| Subject line | 100 | Open rate needs enough volume to clear natural send-time variance |
| Opening line (first sentence) | 150 | Reply rate is lower than open rate so you need more sends to see a signal |
| CTA (call to action) | 150 | Meeting booked rate is the lowest-frequency metric; small samples are misleading |
| PS line | 200 | PS impact on reply rate is small; you need scale to detect it |
| Full email variant (multiple elements) | Do not test | You cannot isolate cause. Split into single-element tests instead. |
Rule of thumb: if you cannot hit minimum sample within 3 weeks at your current send rate, the test is not worth running until your volume grows.
Test Log Table
Recreate this as a spreadsheet. One row per test. Archive completed rows monthly, do not delete them.
| Field | What to Enter |
|---|---|
| Test ID | Sequential number. E.g. T-001, T-002. |
| Date Started | Date first send went out |
| Date Closed | Date you hit minimum sample on both arms |
| Channel | Cold email / LinkedIn DM / LinkedIn InMail |
| ICP Segment | Be specific. E.g. "SaaS ops leaders, 50-200 headcount, UK" |
| Sequence Step | Step 1 / Step 2 / Step 3 etc. |
| Element Tested | Subject line / Opening line / CTA / PS line |
| Hypothesis | One sentence. E.g. "A subject line referencing a named pain will outperform a curiosity subject line for this ICP." |
| Control Copy | Full text of the control. Paste the exact wording. Note character count in brackets. |
| Variant Copy | Full text of the variant. Paste the exact wording. Note character count in brackets. |
| Sends (Control) | Number of sends |
| Sends (Variant) | Number of sends |
| Open Rate (Control) | % opened. Leave blank if testing reply/meeting rate only. |
| Open Rate (Variant) | % opened |
| Reply Rate (Control) | % replied (any reply) |
| Reply Rate (Variant) | % replied (any reply) |
| Positive Reply Rate (Control) | % replied positively (interested, asked for more, booked) |
| Positive Reply Rate (Variant) | % replied positively |
| Meetings Booked (Control) | Raw number |
| Meetings Booked (Variant) | Raw number |
| Confidence Indicator | See guide below. Enter: Low / Medium / High |
| Failure Mode Flag | See list below. Enter the relevant code or "None" |
| Decision | ADOPT / DISCARD / RETEST. See decision rules below. |
| Decision Notes | Mandatory. At least one sentence explaining the decision. |
| Owner | Name of rep or manager who ran and closed the test |
Confidence Indicator Guide
This is a simplified signal, not a formal p-value calculation. If you have a statistician, use proper significance testing. For most sales teams, use this:
Step 1. Calculate the absolute difference in the metric you care about (e.g. reply rate variant minus reply rate control).
Step 2. Check the table below.
| Sample per Arm | Difference Needed for Medium Confidence | Difference Needed for High Confidence |
|---|---|---|
| 100 | 5 percentage points | 8 percentage points |
| 150 | 4 percentage points | 7 percentage points |
| 200 | 3 percentage points | 6 percentage points |
| 300+ | 2 percentage points | 4 percentage points |
Traffic light key for the Confidence Indicator field:
- High: Difference meets High threshold. Safe to make a decision.
- Medium: Difference meets Medium but not High threshold. Decision is directional, note the caveat.
- Low: Difference is below Medium threshold. Do not adopt or discard. Either retest at higher volume or park the question.
Common Failure Mode Codes
Enter the relevant code in the Failure Mode Flag column. If more than one applies, list both.
| Code | Failure Mode | What It Means |
|---|---|---|
| FM-01 | Send-day confound | Control and variant sent on different days of the week or around a bank holiday |
| FM-02 | List quality drop | One arm pulled from a fresher or higher-quality list than the other |
| FM-03 | Sequence position drift | Reps changed which step the email was used in mid-test |
| FM-04 | Multiple elements changed | More than one thing differs between control and variant |
| FM-05 | Volume imbalance | One arm had more than 20% more sends than the other |
| FM-06 | ICP mix drift | Segment definition shifted during the test period |
| FM-07 | External event | Industry news, product launch, or market event that skewed responses in the test window |
If a failure mode flag is present, the Decision field must default to RETEST unless the flag is FM-07 and the event is clearly non-recurring.
Decision Rules
You must complete the Decision field before a test row is considered closed.
ADOPT: Variant beat control, Confidence is High, no failure mode flags. Roll the variant into your live sequence as the new control. Note the date adopted and which sequence it entered.
DISCARD: Control beat variant, or variant showed no meaningful difference, Confidence is High, no failure mode flags. The hypothesis is rejected. Note what you learned and what you will not test again.
RETEST: Confidence is Low or Medium, or a failure mode flag is present. Write a note specifying what needs to change before the retest (larger sample, cleaner list, controlled send days).
Do not leave the Decision field as "pending" after the close date has passed. That is how institutional knowledge dies.
Monthly Summary Row
Add one of these rows at the bottom of each month's tests before you archive the log.
| Field | Enter |
|---|---|
| Month | E.g. June 2025 |
| Total tests run | Count of rows |
| Tests closed with High confidence | Count |
| ADOPTs | Count and list Test IDs |
| DISCARDs | Count and list Test IDs |
| RETESTs carried forward | Count and list Test IDs |
| Net improvement in reply rate vs. month start | Calculate from your baseline and current best-performing variant |
| Key insight for pipeline review | One or two sentences. E.g. "CTA framing around a specific outcome consistently outperforms open-ended asks for this ICP. Two subject line tests were inconclusive at current volume." |
| Presented by | Name |
Quick-Reference: What Goes in the Copy Fields
Be exact. Vague copy notes make the log useless in six months.
Subject line test example:
- Control:
Quick question, [First Name](28 chars) - Variant:
[First Name], are you still losing deals to [Competitor]?(57 chars)
Opening line test example:
- Control:
I work with [role] at [company type] who are struggling to hit [metric]. - Variant:
[First Name], I noticed [Company] recently [specific trigger]. That usually means [pain].
CTA test example:
- Control:
Worth a 20-minute call this week? - Variant:
Are you the right person to speak with about this, or is there someone else on your team I should reach out to?
Paste the full text every time. Do not paraphrase.