Stop signing off AI projects on "it feels much faster"
An acceptance standard anyone can take and use — whether you work with us is beside the point. We publish it because this industry lacks a ruler, and the lack of one hurts everyone who delivers honestly.
DOCUMENT NO.TP-STD-001
VERSIONv1.0 · current
PUBLISHED2026-07
LICENSEPublic · free for anyone to use
FOREWORDWhy a ruler is needed
AI projects rarely die of bad technology. They die because nobody can say what they actually delivered.
Three months in, the CEO asks what was saved. Operations says efficiency improved. Finance says the books show nothing.
IT says call volume is up fivefold. Three answers to three different questions, and not one of them can be signed.
Typical fake acceptance: "40% faster" (than what?) · "95% accurate" (on which test set, written by whom?) · "significant headcount savings" (how many people, and what do they do now?) · "positive user feedback" (how many users, asked what?)
This standard is the ruler: a shared vocabulary, five rules, and a one-page template that is ready to sign.
1Scope
This standard applies to the acceptance of any enterprise AI project whose goal is to replace or assist human work in an existing business process —
any industry, any model, any vendor; in-house teams included.
It does not evaluate model benchmarks. It answers one question only: what did this project actually deliver, and can the number be signed?
2Terms and definitions
2.1Manual baseline
The measured cost of running the process by hand, taken before AI goes live. Four numbers: minutes per item, labor cost per item, monthly volume, rework rate.
2.2Golden set
A set of correct answers labeled by the client's own domain experts (200–500 items recommended, hard cases included), used to score system quality. It belongs to the client and is version-controlled.
2.3Acceptance line
The quality floor both sides fix before work begins, expressed as a golden-set score threshold (e.g. score ≥ 95%). Once written, neither side may change it unilaterally.
2.4Full AI cost
Model calls + platform/license fees + amortized operations effort. All three count; a missing line is a missing line.
2.5Benefit
Labor cost per item × actual AI-handled volume − full AI cost. Volume comes from system records, never estimates; periods below the acceptance line are excluded.
3The rules (five)
RULE 3.1
Measure the manual baseline before touching AI
Before anything goes live, measure what the workflow costs today — the four numbers in 2.1. Without that anchor, every later "improvement" floats free.
The baseline must be measured before work starts. A baseline reconstructed afterwards always drifts in whichever direction suits the person reconstructing it.
RULE 3.2
Both sides sign the baseline, with names and dates
A baseline the vendor reports alone does not count. Neither does one the client sets alone. The signature is not ceremony — it closes the door on disowning the number later. Three things go on file: signatories, date, and the definition used.
RULE 3.3
Score quality on a golden set, not with adjectives
Quality is the system's score on the golden set (2.2). One number, no adjectives.
The vendor may read the items and tune against them, but may not change the answers.
RULE 3.4
Set the acceptance line before you see results
The acceptance line (2.3) is fixed before work begins. A line drawn after seeing the results is not a line. Missing it means missing it: remediate until it passes, or end the work — the one thing you may not do is move the line.
RULE 3.5
Benefit = manual baseline × actual volume − full AI cost
Every variable has a source: the baseline from the signed record (3.2), the volume from system records, the cost with all three lines of 2.4 — never model calls alone.
A workflow below the acceptance line contributes nothing to the benefit figure. This rule is the easiest to sidestep and the most important.
4One-page template
Fill it in as is — no tooling required. Whichever field you cannot fill is where the project's risk lives.
[ACCEPTANCE BASELINE — SIGNED RECORD]
Workflow name: ____________________
Scope: ____________________ (what counts, what does not)
1. Manual baseline (measured before AI goes live)
Minutes per item: ______ Measured how: ____________________
Labor cost per item: ______ Basis: ______ salary / ______ productive hours
Monthly volume: ______ Source: ____________________
Current rework / error rate: ______ %
2. Quality standard
Golden set size: ______ Labeled by: ____________________
Golden set version: ______ Stored at: ____________________
Acceptance line: score >= ______ % (fill in before work starts; not editable afterwards)
3. Benefit definition
Benefit = labor cost per item x actual AI-handled volume - full AI cost
Full AI cost includes: [ ] model calls [ ] platform / license [ ] operations effort
Periods below the acceptance line: [ ] excluded from benefit
4. Signatures
Client: ____________ Title: __________ Date: __________
Vendor: ____________ Title: __________ Date: __________
5Known evasions (recognize them early)
Inflated baseline — costing the work at the most senior person's rate, or folding in hours that belong to other processes.
Golden-set drift — the vendor "helps improve" the labels, and the answers shift toward what the system is good at.
Scope drift — hard cases quietly move outside the "applicable scope" after launch, and the score rises on its own.
Missing cost lines — counting model calls but not platform fees or operations effort.
Cherry-picked windows — reporting only the best two weeks.
APPENDIXLicense and versions
License: this standard is published openly. Anyone may use, copy, modify, and apply it commercially — no attribution, no need to tell us. Take it into a negotiation with another vendor if you like: if it saves you from one muddled project, it has done its job.
Why we publish it: our business is making the numbers add up. The more people who count honestly, the easier our work gets; the more people who only tell stories, the harder it gets for everyone.
v1.0 · Published 2026-07 · Suggestions to [email protected] — contributors are credited in the next revision.
Want someone alongside you for the first baseline?
That is exactly what week one of a Bootcamp is — pricing takes three sentences, all on the Bootcamp page