Agencies

AI quality control for client deliverables, before a human reads them

Every deliverable checked before a human sees it. The system writes and grades work against your own standards and your client's tone, names the exact passage that fails, and holds back anything below your passing mark. Editing hours collapse into minutes.

Agencies: AI quality control for client deliverables, before a human reads them

Why agency review time grows faster than the agency

Quality in an agency lives in a few heads. A senior person knows what goes out and what does not. That knowledge is real, it is just not written down anywhere, so the only way to apply it is to read every piece end to end.

That works at five deliverables a week. At fifty it becomes the bottleneck. Review gets pushed to the evening, then it gets skimmed, then something goes out that should not have. The correction costs more than the review would have, and it costs it in front of the client.

The second problem is drift. Every client has a tone. Formal or informal, first person or third, claims allowed or claims banned. New writers learn it slowly, freelancers never quite learn it, and nobody notices the slide until the client does.

  • The same two mistakes come back every month, from different people.
  • Only one or two people in the company can sign off, so everything queues behind them.
  • Feedback is written from scratch each time instead of pointing at a rule.
  • A piece reaches the client with the wrong form of address for that account.
  • Numbers and legal claims get copied forward from an old draft, unchecked.
  • Nobody can say which criteria a piece actually passed. Only that it felt fine.

How an automated quality gate works on real deliverables

The gate sits between the person who produced the work and the person who signs it off. Nothing reaches a client without passing through it, and nothing passes because it looks confident.

  1. 01

    Your standard becomes a written list

    We turn the knowledge in your seniors' heads into named criteria: claims backed by a source, structure matches the brief, no filler adjectives, call to action present and specific, legal limits for that industry. Each one gets a definition of what full marks look like and what a fail looks like. This is the part most teams never get to on their own, and it is the part that makes everything after it possible.

  2. 02

    Each client gets their own tone profile

    The formal client and the casual client are graded differently on the same criterion. The profile is built from work that account has already accepted, not from a generic style guide. When an account changes its preference, you change it in one place and every future deliverable for that account is graded against the new one.

  3. 03

    Every deliverable is graded, not spot checked

    The gate scores each criterion, quotes the passage in the work that earned the score, and writes the fix. Not a rewritten paragraph: the sentence to cut, the missing element to add, the claim that needs a source. The reviewer reads the two findings instead of the eight hundred words.

  4. 04

    Facts get verified against the source, not the draft

    Where a piece states a figure, a deadline or a legal rule, the system checks it against the source document rather than trusting the previous draft. A number that cannot be traced to a source is flagged as unverified. It is never smoothed over.

  5. 05

    Below the passing mark, it stops

    You set the mark. Work above it moves on with its findings attached. Work below it is held back and named, and the failing criteria are listed. The system does not decide that a borderline piece is probably fine. Anything ambiguous, and anything touching money, promises or legal exposure, waits for a person by design.

  6. 06

    Every correction feeds the standard

    When a reviewer overrides a score or adds a note, that becomes evidence. The criteria sharpen, the client profile gets more accurate, and the same argument does not have to be had twice. Over months the gate stops being a filter you installed and becomes the written form of how your agency works.

Take this with you

A grading prompt that judges work and refuses to rewrite it

This is the single handgrip at the centre of the whole thing: grade one piece against one written standard. Paste the work, paste your criteria, get a score per criterion with the passage that proves it and one concrete change. It is deliberately forbidden from rewriting the piece for you, because a model that rewrites hides the mistake instead of teaching it. It is also forbidden from guessing when the material gives it nothing to judge on.

prompt
ROLE
You grade one piece of work against a fixed standard.
You are a reviewer, not a writer. You never produce an improved version.

INPUT
[STANDARD]
  <one criterion per line. For each: the name, one line on what a 10 looks
   like, one line on what a 4 looks like.>
[BRIEF]
  <what was asked for: audience, format, length, required elements, tone.>
[WORK]
  <the finished piece, in full.>
[PASSING MARK]
  <the total score at or above which this piece may go out.>

RULES
1. Grade only the criteria in [STANDARD]. If something else in [WORK] looks
   wrong to you, stay silent about it. The standard is the standard.
2. For every criterion, give three things:
   a score out of 10,
   one passage quoted word for word from [WORK] that justifies the score,
   one concrete change, naming the sentence and saying what to do with it.
3. Quote verbatim. Never paraphrase a quote, never invent one, never quote
   from [BRIEF] as if it were [WORK].
4. A change is an instruction, not a replacement text. "Cut the adjective
   in sentence four" is allowed. Handing me the rewritten sentence is not.
5. If [WORK] and [BRIEF] give you no basis to judge a criterion, write
   "cannot judge: no basis in the material" and give no score for it.
   Do not average, do not assume, do not be generous.
6. No praise, no encouragement, no summary of what the piece does well
   unless a criterion asks for it.

OUTPUT
A table, one row per criterion:
  Criterion | Score | Evidence (quoted) | Change (one sentence)
Then, in this order:
  Total: <sum> out of <maximum possible>
  Verdict: PASS, or HOLD with the failing criteria named
  Cannot judge: <list, or "none">
  The two changes that move the score most, in order.
  1. Build your standard from evidence, not opinion. Put three pieces you accepted and three you sent back on the table, and for each rejected piece write the one sentence reason it came back. Those reasons are your criteria. Six to eight is enough. Fewer than four and the grade is meaningless.
  2. Write each criterion twice: what a 10 looks like, what a 4 looks like. If you cannot describe the 4, the criterion is not a criterion yet, it is a preference.
  3. Set a passing mark and write it down before you grade anything. Setting it afterwards means setting it to whatever the piece scored.
  4. Run it on the three pieces you already accepted. If they fail, your standard is wrong, not the work. Fix the wording and run it again. This calibration pass takes twenty minutes and is the difference between a useful grader and a noisy one.
  5. Keep one saved version of the standard per client, with that client's tone written into the tone criterion in their own words.

It grades one piece that a person pastes in. It does not know your client, so you carry the tone in by hand every time. It cannot check a figure against a source document, so a wrong number that reads confidently will pass. It forgets everything the moment the window closes, which means last month's corrections do not sharpen next month's grade. And it only runs when somebody remembers to run it, which on a busy week is exactly when nobody does.

From a prompt you run to a gate nothing gets past

The prompt above is the check. The built system is the gate. The difference is not intelligence, it is position: the gate sits in the path of the work, so no deliverable can route around it on a Friday afternoon.

It also knows things a pasted prompt cannot know. Which client this is for and how that client wants to be spoken to. What the brief actually asked for, read from the brief itself rather than retyped. Whether the figure in paragraph two matches the source it came from. What your reviewers overruled last month and why.

And it is honest about its own edges. Ambiguous cases are handed to a person rather than resolved by confidence. Anything involving a promise, a price or a legal claim goes to a human by default, no matter how high it scored. The value of a gate is that you can trust the passes, and you can only trust the passes if the system is willing to say it does not know.

  • Every deliverable is graded automatically, before the first human reading.
  • Criteria and tone are held per client, not retyped per piece.
  • Facts are checked against the source document, not against the previous draft.
  • Work below the passing mark is held back and named, not quietly passed.
  • Reviewer overrides feed back into the standard, so the same fight happens once.
  • Anything ambiguous, or touching money or promises, waits for a person on purpose.

Frequently asked questions

Can I not just do this with ChatGPT?

For one piece at a time, yes, and the prompt on this page is exactly that. The limits show up at scale. A chat window does not know which client the work is for, cannot read the brief or the source document by itself, and forgets every correction you made last week. It also only runs when a person remembers to open it, so the deliverable that most needs checking is the one that goes out unchecked. The built system sits in the path of the work instead of waiting to be asked.

Will an AI reviewer flatten our writing into generic text?

Not if it is forbidden from writing. The gate grades and points, it does not rewrite. Your people make the change, in your voice. That is a deliberate design choice: a system that silently improves text teaches nobody anything and slowly replaces your style with its own.

What happens to a deliverable that fails the check?

It is held back and named, with the failing criteria and the exact passages listed. It does not go to the client. The author gets the findings, fixes them, and the piece is graded again. You set the passing mark, so you decide how strict the gate is, and you can set it differently per client.

How do we define our quality standard if we have never written it down?

You do not need a document to start. Take three pieces you accepted and three you sent back, and write the one line reason each rejection came back. Those reasons are the first criteria. We build the standard with your seniors from real work rather than from a blank page, and it sharpens every time a reviewer overrides a score.

How much does a quality gate like this cost to build and run?

It depends on how many criteria you check, how many clients have their own tone, and what has to be verified against source documents. The first version is deliberately small: one team, one deliverable type, your existing standard. We scope it against your real work before quoting, and you see the shape of the system before you commit to it.

Where does our work and our clients' data live?

In your environment, under your accounts. The system reads the deliverables and briefs it needs and writes its findings back where your team already works. We do not build systems that require handing your client material to a third party you have no contract with, and we tell you exactly which services touch which data before anything is built.

Does this replace our editors?

It replaces the part of their job that is reading eight hundred words to find two things. They still make the judgment calls, still own the client relationship, and still decide what good means. What changes is that they read the findings instead of the full piece, and that their decisions get written into the standard instead of staying in their heads.

Put a gate in front of your delivery

If your review queue sits behind one or two people, or the same corrections keep coming back, that is a system we can build. Apply and we will look at your actual deliverables and your actual standard first.

Apply

Related use cases