AI can score a person in seconds. What it can't do is verify itself. Second Marking is a calibrated panel of trained human assessors who mark a sample of your conversations blind, show you exactly where they agree and diverge from the AI, and give you a published reliability figure you can put in front of your own clients.
More and more organisations let AI judge how people perform — roleplay tools, call-scoring engines, interview sifts, quality assurance, capability platforms. It's fast and it scales. But when a client asks "how do we know these scores are right?", the honest answer is: the same system that produced them can't be the one that checks them. Neither can the party whose programme is being assessed.
An AI vendor grades and reports on its own output. There is no second, disinterested read to confirm it — so the number is trusted, not verified.
These grades end up in performance conversations, capability reports, and client deliverables. If one is challenged, someone needs to be able to stand behind it.
A trained assessor, working blind, marking a sample against the same rubric — and a measured agreement figure showing where the AI can be relied on and where it can't.
You decide what share of each client's conversations gets an independent human mark — a spot check, half, or all of them. We handle the rest.
Recorded roleplays or real calls, with the AI's scores alongside. Everything is anonymised before an assessor sees it — see Data & privacy.
Trained assessors mark against your own rubric, blind — no sight of the AI's score or the person's name. One in five is marked twice, by two assessors independently.
Per criterion, we show where the human marks and the AI agree, where they diverge, and by how much — with the evidence clip behind every judgement.
Each quarter, a signed statement of how consistently the panel agreed — something you can put in front of your own clients as proof the scoring has been independently checked.
It isn't only sales. Wherever a model judges a person and someone acts on the result — coaching, hiring, promotion, performance — the same question applies: has anyone checked?
You put your name on scoring your clients did not build and cannot inspect. Second Marking is the independent check that lets you tell a client the scoring has been stress-tested by people, not just an AI.
Your product scores conversations at scale. An independent human agreement figure — and the data-quality catches a score can't make — is the assurance layer that makes buyers trust the number.
Add an evidence-backed human mark to your programme without hiring and calibrating an assessment team yourself. We are the reserved, standing panel.
AI now grades a share of your interactions that no QA team could ever listen to. That reach is the point, and it is also why an unchecked scorecard quietly sets the standard for hundreds of agents at once.
Once a score influences who gets coached, promoted or performance-managed, someone will be asked to justify it. Better to hold the agreement figure before the question arrives than after.
Where a capability score about a named individual has to survive being questioned, the clip, the reasoning, the rubric version and the marker reference all travel with it.
This section answers the questions a client's data or compliance team will ask before a single conversation is assessed. It's written to be shared — send it on, or point a prospect straight here.
We assess the conversations you send us — we don't run the roleplays or record the calls ourselves. So the answer depends on what you send, and in both cases it's designed so there's nothing sensitive to expose.
Roleplays. We're often asked to assess roleplays against an AI buyer. There, the person is talking to a simulated counterparty, not a real customer — so there is no genuine third party on the line and no information about your clients' customers is present at all. The assessment is of how the person handled the conversation — the skill, not the deal.
Real sales calls. Where you ask us to assess real commercial conversations, which we also do, that's where our protocols come in: full anonymisation of the people involved, and — where the content warrants it — a sanitisation step that strips business and personal detail before an assessor sees anything. Both are covered below.
No. Every person is given a code — for example P-07. The assessor sees only that code, the anonymised transcript, and the AI's scores. They never see a name, an email address, a job title, or any other identifying detail.
Personal details can occasionally surface in the course of a conversation — a first name spoken aloud, say — but they are never the focus, are not surfaced deliberately, and results are only ever reported against the code. The key that maps a code back to a person is held by you, not by us, so on our side the marking cannot be tied to a named individual at all.
Before a conversation reaches an assessor, it can be run through a series of AI agents that sanitise and anonymise it: stripping business context, pricing, product detail, and personal information such as names, phone numbers, addresses and email addresses. Only the cleaned version is delivered for marking. For real-conversation work this is the default; for roleplays, where there's no real party on the line, it's applied where a client wants the extra assurance.
Two things worth being explicit about. First, those agents run on models we host ourselves — sanitisation happens inside our own infrastructure and your data is not sent to any third-party AI service to be cleaned (more on that below). Second, we're open about the trade-off: the sanitising agents can be over-zealous — occasionally removing context that a grade actually depends on, or cutting a conversation in a way that makes some criteria harder to assess. Because business or personal detail sometimes ties directly to a judgement, heavy sanitising can very slightly reduce assessment accuracy. It's a deliberate exchange of a little precision for a lot of privacy, applied where it's warranted rather than blanket.
No, on both counts. This is where a lot of assurance offers quietly fall down, so we're specific about it. The AI agents that sanitise and anonymise conversations run on models we host on our own infrastructure — your data is not sent to OpenAI, Anthropic, or any other third-party AI service to be processed. There is no external API holding a copy, and nothing you send us is ever used to train a model, ours or anyone else's.
The only AI scoring in the picture is the one you supply alongside the conversation — the output we're independently checking. Our role adds a human panel and, where used, in-house sanitisation. No part of the pipeline hands your material to an outside model.
Access is limited to a small, named panel of trained assessors, each under a written confidentiality agreement, on least-privilege access that is authenticated and logged. Assessors work through an anonymised interface and cannot export raw material.
Each client's data is held in its own separate silo — separate storage and separate export, never pooled with, or visible to, another client. We use your data solely to perform the assessment you've asked for: never to benchmark other clients, populate a shared library, or any purpose beyond the engagement.
Conversations and results are held on established cloud infrastructure with encryption in transit and at rest. Access is authenticated, logged, and limited to the people who need it. Data is retained only as long as needed to deliver and stand behind the assessment, and is deleted on request, with a deletion certificate provided. Retention periods can be fixed in the contract.
Hosting-region and residency requirements are scoped per engagement: tell us where your data must and must not sit, and we'll confirm what we can commit to in writing before a single conversation is processed, rather than claim a blanket guarantee that wouldn't hold for every case.
Yes. You are the data controller and we act as your data processor under a written Data Processing Agreement — the data is yours, processed only on your documented instructions and only to perform the assessment. Our standard DPA covers purpose limitation, confidentiality, security measures, sub-processor terms, assistance with data-subject requests, international transfers, breach handling and deletion. Read the full DPA →
We maintain a current list of sub-processors and share it on request, with notice of any change. Where a processing activity or transfer requires them, standard contractual clauses (or the UK IDTA) are included. If your legal team would rather work from your own paper, send it over — we'll review and sign from there.
Controls in place today: encryption in transit and at rest; authenticated, logged, least-privilege access; per-client data isolation; self-hosted AI so material never leaves our environment for a third-party model; and a confidentiality agreement on every assessor.
On formal certification, we'll be straight with you. We don't yet hold SOC 2 or ISO 27001. Our controls are modelled on the ISO 27001 framework and formal certification is on our roadmap. In the meantime we're glad to complete your security questionnaire, share our controls documentation, and support a due-diligence review — we'd rather show you where we actually are than display a badge we haven't earned.
If an incident affecting your data occurred, we would contain it, investigate, and notify you without undue delay — with what we know, what's affected, and what we're doing about it — and support any notifications you in turn need to make to your own clients or to a regulator. The specific notification window — 72 hours — and the responsibilities on each side are set out in the DPA, so they're contractual, not just a promise on a page.
Yes. The panel is a standing, calibrated resource rather than people assembled per project, so volume scales without recalibrating from scratch each time. We agree throughput and turnaround with you up front and hold capacity against a committed volume, with a service level set in writing for large or ongoing programmes.
Continuity is built into the method: because a sample of every batch is double-marked and every assessor is calibrated to a common standard, no single marker is a point of failure — work re-routes within the panel without changing how a score is reached. The reliability figure we publish each quarter is also the early warning if consistency ever drifts as volume grows.
The independence is enforced in the system, not by policy. An assessor never sees the AI's score, the vendor's report, another assessor's marks, or a real name — because the moment a marker is anchored to the answer they're supposed to be checking, the comparison is worthless.
One conversation in five is independently double-marked by a second assessor, and the inter-rater agreement is measured and published each quarter. That figure is what turns a mark from an opinion into a measurement — and it's reported in full, not just when it flatters us.
Independent human assessment by a calibrated standing panel. Tell us what you're scoring and we'll show you what an assured version looks like.