An AI draft can read smoothly while missing the point of the brief. A report summary can repeat the right figures and still draw an unsupported conclusion. Before a team relies on a workflow, it needs a shared way to decide whether the result is usable.
That does not have to begin with a complicated scoring system. Start with examples of the task and the questions a good reviewer would ask. The marketing scenarios below are illustrative.
Use examples that expose the difficult parts
Gather approved examples of the work you want the assistant to handle. Include an ordinary request, an incomplete brief, conflicting information and a case where the correct response is to ask a question. Remove information the proposed tool is not approved to receive.
For each example, write down the source material and the expected behaviour. Keep some examples aside while adjusting the instructions, then use them to check whether the improvement carries over to work the assistant has not been tuned around.
Separate accuracy from writing quality
A reviewer needs to see both. For a campaign draft, check the product claims against the approved brief before judging tone. For a summary, check what was left out as well as what was included. Make the review criteria specific to the task.
- Facts: are names, dates, numbers and product details supported by the source?
- Completeness: did the output retain the request and the important constraints?
- Uncertainty: does it identify missing or conflicting information?
- Usefulness: can the next person act on it without reconstructing the task?
- Voice: would your audience understand it, and does it sound like your organisation?
Mark serious errors separately from minor edits. A misleading product claim should not disappear inside a favourable average score for grammar, format and tone. Record what needed correction so the next review uses the same standard.
Check the action as well as the answer
Anthropic’s January 2026 guide, Demystifying evals for AI agents, distinguishes what an agent says it did from the actual outcome. It also describes combining automated checks with human judgement.
For a marketing handoff, check that the right record reached the right person with the right information. A message saying “done” is not enough. During a pilot, use an agreed test destination and inspect the result before enabling actions that affect customers.
For a writing task, the outcome includes the work a human has to do afterwards. Record review time, the corrections required and whether the draft was used. These observations are more useful than asking the assistant to rate its own answer.
Keep a small set of repeatable checks
Save the examples, source versions, instructions and review notes together. When the workflow changes, run the same checks again. Add new examples when the team encounters a problem the current set does not cover.
If you use AI to help review a large volume of drafts, compare its decisions with an experienced human reviewer on the same examples. Discuss disagreements before treating an automated score as an approval. The person responsible for the content should understand what the score measures.
Decide who can approve the result
Agree which outputs stay as drafts, which actions need approval and who owns the review process. For an agency, this may involve both the agency team and a client’s approver. Put that decision into the workflow so it remains clear after the initial pilot.
You are ready to expand when the team can explain what it has checked, which errors still occur and how those errors are handled. The review process should make that decision easier to justify.
We help teams practise this on their own work through AI training.



