Our evaluation method
Agent GPA: we do not ship an agent until it outscores the process it replaces
That sounds obvious. Almost nobody does it. The usual sequence is build the agent, demo the agent, deploy it, and discover six weeks later that the team quietly went back to the spreadsheet. Scoring first is what prevents that, and it is why we sometimes tell clients not to build at all.
First we score the humans
Before writing any agent, we measure the manual process on four things. Not to make a case for automation — to find out whether there is one.
Time per case
How long one person takes on one instance of the task, measured across enough cases to be honest about the slow ones rather than the demo one.
Error rate
How often the manual process gets it wrong today. This number is almost always worse than anyone expects, and it is the single most useful thing scoring produces.
Correction cost
What each error costs downstream — the rework, the senior review time, the customer call. An error that takes four minutes to catch and an hour to unwind are not the same error.
Escalation quality
What happens on the cases the process cannot handle. A clean handoff to a person is a good outcome. A silent wrong answer is not.
Then every version of the agent is scored on the same cases
The same case set, the same four measures, version after version. That is the whole idea, and the discipline is in using the same cases rather than the ones the agent happens to do well on.
An agent goes live when it beats the baseline on the measures that matter for that workflow — and not before. Sometimes that is version two. Sometimes it is version nine. Occasionally it is never, and that is a result rather than a failure.
Sometimes the answer is do not build it
If a process is already 99% accurate and takes four minutes, an agent is not the win. We would rather say so at scoping than eleven weeks in.
The baseline is usually worse than anyone thinks
Manual three-way invoice matching does not run at 100%. Once you actually measure it, the bar the agent has to clear becomes honest.
Regulated clients need this on paper
Under Bank Negara or MAS scrutiny, “the vendor said it works well” is not an answer. A scored evaluation is the artefact your risk function asks for.
Worked examples of agents built this way are in our agent use-case library, and the build work itself is described on AI agent development.
Questions about Agent GPA
Agent GPA is how we decide whether an AI agent is worth shipping. Before any agent is built, we score the existing manual process on four measures: time per case, error rate, correction cost and escalation quality. That becomes the baseline. Every version of the agent is then scored against the same cases, and it does not go live until it outscores the process it would replace. The name is ours; the practice is ordinary evaluation discipline applied earlier than most vendors apply it.
Because without a baseline, 'the agent is 92% accurate' means nothing. Ninety-two per cent is excellent if your team runs at 80% and unacceptable if they run at 99%. Scoring the manual process first is also the step that most often kills a project — and killing it at scoping costs a scoping conversation, while killing it after launch costs a quarter.
Often enough that we consider it a feature of the method rather than a failure of it. If a process is already 99% accurate and takes four minutes, an agent is not the win — the money is better spent elsewhere. The same applies when the data feeding the process is wrong: connecting an agent to bad source data produces confidently wrong answers faster.
Yes, and it is usually the reason they ask. If you are under Bank Negara Malaysia or MAS scrutiny, 'the vendor said it works well' is not an answer. A scored evaluation with a documented baseline, a per-version record, and a named human accountable for sign-off is the artefact your risk function will ask for. We produce it as a matter of course rather than on request.
The Claude Certified Architect who designed the system. Not a delivery manager relaying numbers, and not the client discovering them after go-live. You approve the baseline before we build and the scores before anything touches a customer or a ledger.
Free consultation
Tell us the process you want scored
We will measure what it costs you today before proposing anything. If the numbers say an agent is not the answer, that is what we will tell you.
