ThinkWork
Team/Manager Template Free

Cold Outreach A/B Test Log & Decision Tracker

A structured template for logging every cold email and LinkedIn message A/B test your team runs, complete with copy, sample sizes, results, confidence indicators, and a mandatory go/no-go verdict. Builds a proprietary, compounding record of what works for your specific ICP over time.

How to use it

Open a new spreadsheet and recreate the tables below as tabs, or paste them directly into your existing sales ops tracker. Log every test before you send it, fill in results once you hit minimum sample size, and complete the Decision field immediately so findings don't get lost. Review the Monthly Summary rows in your pipeline meeting each month.

What's inside

  • Pre-test setup fields covering hypothesis, element tested, channel, and ICP segment
  • Side-by-side control vs. variant copy log with character counts
  • Minimum sample size guide by test type (subject line, opening line, CTA, PS line)
  • Results table with open rate, reply rate, positive reply rate, and meeting booked rate per arm
  • Statistical confidence indicator with a simple calculation guide and traffic-light key
  • Mandatory Decision field with three options: Adopt, Discard, or Retest, each requiring written notes
  • Common failure modes column with named causes such as send-day confound and list quality drop
  • Monthly Summary row for manager reporting in pipeline reviews
  • When NOT to run an A/B test guidance note

Purpose

Every team has opinions about subject lines. Few have evidence. This tracker turns your outreach experiments into a proprietary record, so the knowledge stays in the team when reps leave, compounds over time, and gives your manager something real to show in pipeline reviews.

This is not a grading tool for a single email. It is not a benchmark comparison. It is a running log of controlled experiments, each with a clear verdict.


When NOT to run an A/B test

Before you open a new row, check these conditions. If any apply, fix them first.

  • Your send volume is under 50 per week per arm. You will not reach minimum sample in a reasonable timeframe and the result will be noise.
  • Your list quality has changed between sends (different source, different scrape date, different ICP tier). The test is confounded before it starts.
  • You changed more than one element between control and variant. You will not know what caused the difference.
  • You are mid-sequence. Test on step 1 first. Multi-touch sequences amplify noise from early steps.

Minimum Sample Size Guide

Use this before you start a test, not after. These are floors, not targets. Aim higher if you can.

Element TestedMinimum per ArmRationale
Subject line100Open rate needs enough volume to clear natural send-time variance
Opening line (first sentence)150Reply rate is lower than open rate so you need more sends to see a signal
CTA (call to action)150Meeting booked rate is the lowest-frequency metric; small samples are misleading
PS line200PS impact on reply rate is small; you need scale to detect it
Full email variant (multiple elements)Do not testYou cannot isolate cause. Split into single-element tests instead.

Rule of thumb: if you cannot hit minimum sample within 3 weeks at your current send rate, the test is not worth running until your volume grows.


Test Log Table

Recreate this as a spreadsheet. One row per test. Archive completed rows monthly, do not delete them.

FieldWhat to Enter
Test IDSequential number. E.g. T-001, T-002.
Date StartedDate first send went out
Date ClosedDate you hit minimum sample on both arms
ChannelCold email / LinkedIn DM / LinkedIn InMail
ICP SegmentBe specific. E.g. "SaaS ops leaders, 50-200 headcount, UK"
Sequence StepStep 1 / Step 2 / Step 3 etc.
Element TestedSubject line / Opening line / CTA / PS line
HypothesisOne sentence. E.g. "A subject line referencing a named pain will outperform a curiosity subject line for this ICP."
Control CopyFull text of the control. Paste the exact wording. Note character count in brackets.
Variant CopyFull text of the variant. Paste the exact wording. Note character count in brackets.
Sends (Control)Number of sends
Sends (Variant)Number of sends
Open Rate (Control)% opened. Leave blank if testing reply/meeting rate only.
Open Rate (Variant)% opened
Reply Rate (Control)% replied (any reply)
Reply Rate (Variant)% replied (any reply)
Positive Reply Rate (Control)% replied positively (interested, asked for more, booked)
Positive Reply Rate (Variant)% replied positively
Meetings Booked (Control)Raw number
Meetings Booked (Variant)Raw number
Confidence IndicatorSee guide below. Enter: Low / Medium / High
Failure Mode FlagSee list below. Enter the relevant code or "None"
DecisionADOPT / DISCARD / RETEST. See decision rules below.
Decision NotesMandatory. At least one sentence explaining the decision.
OwnerName of rep or manager who ran and closed the test

Confidence Indicator Guide

This is a simplified signal, not a formal p-value calculation. If you have a statistician, use proper significance testing. For most sales teams, use this:

Step 1. Calculate the absolute difference in the metric you care about (e.g. reply rate variant minus reply rate control).

Step 2. Check the table below.

Sample per ArmDifference Needed for Medium ConfidenceDifference Needed for High Confidence
1005 percentage points8 percentage points
1504 percentage points7 percentage points
2003 percentage points6 percentage points
300+2 percentage points4 percentage points

Traffic light key for the Confidence Indicator field:

  • High: Difference meets High threshold. Safe to make a decision.
  • Medium: Difference meets Medium but not High threshold. Decision is directional, note the caveat.
  • Low: Difference is below Medium threshold. Do not adopt or discard. Either retest at higher volume or park the question.

Common Failure Mode Codes

Enter the relevant code in the Failure Mode Flag column. If more than one applies, list both.

CodeFailure ModeWhat It Means
FM-01Send-day confoundControl and variant sent on different days of the week or around a bank holiday
FM-02List quality dropOne arm pulled from a fresher or higher-quality list than the other
FM-03Sequence position driftReps changed which step the email was used in mid-test
FM-04Multiple elements changedMore than one thing differs between control and variant
FM-05Volume imbalanceOne arm had more than 20% more sends than the other
FM-06ICP mix driftSegment definition shifted during the test period
FM-07External eventIndustry news, product launch, or market event that skewed responses in the test window

If a failure mode flag is present, the Decision field must default to RETEST unless the flag is FM-07 and the event is clearly non-recurring.


Decision Rules

You must complete the Decision field before a test row is considered closed.

ADOPT: Variant beat control, Confidence is High, no failure mode flags. Roll the variant into your live sequence as the new control. Note the date adopted and which sequence it entered.

DISCARD: Control beat variant, or variant showed no meaningful difference, Confidence is High, no failure mode flags. The hypothesis is rejected. Note what you learned and what you will not test again.

RETEST: Confidence is Low or Medium, or a failure mode flag is present. Write a note specifying what needs to change before the retest (larger sample, cleaner list, controlled send days).

Do not leave the Decision field as "pending" after the close date has passed. That is how institutional knowledge dies.


Monthly Summary Row

Add one of these rows at the bottom of each month's tests before you archive the log.

FieldEnter
MonthE.g. June 2025
Total tests runCount of rows
Tests closed with High confidenceCount
ADOPTsCount and list Test IDs
DISCARDsCount and list Test IDs
RETESTs carried forwardCount and list Test IDs
Net improvement in reply rate vs. month startCalculate from your baseline and current best-performing variant
Key insight for pipeline reviewOne or two sentences. E.g. "CTA framing around a specific outcome consistently outperforms open-ended asks for this ICP. Two subject line tests were inconclusive at current volume."
Presented byName

Quick-Reference: What Goes in the Copy Fields

Be exact. Vague copy notes make the log useless in six months.

Subject line test example:

  • Control: Quick question, [First Name] (28 chars)
  • Variant: [First Name], are you still losing deals to [Competitor]? (57 chars)

Opening line test example:

  • Control: I work with [role] at [company type] who are struggling to hit [metric].
  • Variant: [First Name], I noticed [Company] recently [specific trigger]. That usually means [pain].

CTA test example:

  • Control: Worth a 20-minute call this week?
  • Variant: Are you the right person to speak with about this, or is there someone else on your team I should reach out to?

Paste the full text every time. Do not paraphrase.

Stay current

Get told when new Cold Email, LinkedIn & Social Selling resources land.

Tick the topics you care about, then subscribe. Alerts start straight away, no confirmation step. No digest spam, just a note when something genuinely useful is added.

Choose topics below (or leave blank for everything). Unsubscribe any time.