surehand
All articlesOperations

Confidence thresholds for AI agents: what the score means, and why it drifts

Reference5 min readSurehand

Many agents hold a case when their confidence falls below a threshold. "Anything under 80% goes to a person." It sounds like a control. It is one, if 80% means 80%. If the agent says 90% and is right half the time, your threshold lets through the cases it should hold. The property you need has a name: calibration.

In one sentence: a confidence threshold is only as good as the calibration behind it, and you have to measure that on your own cases.

Where it comes from

Statistics and machine learning have measured calibration for years. The scikit-learn guide gives the plain version. A well calibrated classifier's scores "can be directly interpreted as a confidence level"1. Among cases scored close to 0.8, "approximately 80% actually belong to the positive class"1.

Then models got bigger and the property slipped. In 2017, Guo and colleagues found that "modern neural networks, unlike those from a decade ago, are poorly calibrated"2. Language models brought it back into view. Anthropic researchers reported in 2022 that "larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format"4. OpenAI's GPT-4 report found the pre-trained model was well calibrated. But after the post-training that makes it a helpful assistant, the chart caption reads: "The post-training hurts calibration significantly"3. The same report warns GPT-4 "can also be confidently wrong in its predictions"3.

Note the conditions: the right format, the pre-trained model. The agent you buy is neither.

How it works

A confidence score in an agent comes from one of three places. The model's own token probabilities. A number the model writes when you ask it how sure it is. Or a separate check, such as whether the invoice fields matched the PO. Each behaves differently. Ask which one you are looking at.

To set a threshold you can defend, use your own labelled cases. Shadow mode produces exactly these. Here is the method, with illustrative numbers.

Score bandCasesAgent rightReal accuracy
0.95 and up1,2001,18899%
0.90 to 0.9540036892%
0.80 to 0.9025019076%
Below 0.801508456%

Read it by band. In this example the agent is about right at the top and overconfident in the middle: it says 80 to 90 and delivers 76. Now pick the threshold from the error rate you can live with, not from the label. If one mistake in a hundred is your limit for autopay, the line sits at 0.95, not 0.80. Everything below goes to your approver.

Then check the cost of that choice. In this example, 800 of 2,000 cases fall below 0.95. That is 40% of the queue for one person. You may decide to use 0.90 plus hard rules, or split the work differently. The table makes that a decision rather than a guess.

Why thresholds drift

A threshold that was right in March can be wrong in June. Three things move it.

A new model. Different model, different score distribution. The same 0.9 means something else.

A new prompt or rule. Change the instructions and you change what the model is confident about.

New kinds of work. A supplier with a new invoice layout. A claim type you did not see in the shadow period. Scores on unfamiliar cases are the least trustworthy of all.

NIST's AI RMF asks for this kind of check. It lists "valid and reliable" first among the traits of trustworthy AI5. Its Measure function covers tracking over time. In practice: re-run the band table after every model or rule change, and monthly on a sample of live cases.

What it is good at, and what it is not

Good at. Sending the ambiguous middle to a person. Most of the queue is clear. The cases the agent finds hard are often the ones a person should see.

Not good at. Catching confident errors. Prompt injection, a forged invoice, a changed bank detail: the agent can be very sure and very wrong. So confidence must never be the only trigger. Hard facts (amount over a limit, bank details first seen today) should hold a case whatever the score.

Easy to fake. A number the model writes about itself is text, not a measurement. It may be useful. Treat it as unproven until your band table says otherwise.

What to check

  1. /01

    Where does the score come from? Token probability, self-reported, or a separate check?

  2. /02

    Show me the band table on my cases. Not the vendor's benchmark.

  3. /03

    What error rate does the threshold give above the line?

  4. /04

    How many cases fall below it per day, and who gets them?

  5. /05

    Which hard rules hold a case regardless of confidence?

  6. /06

    When was the table last re-run? What changed since?

Where it is going

Our view: confidence will move from the model to the check. Instead of asking the model how sure it is, systems will score what can be verified: did the fields match, did the supplier exist, did the total add up. Those scores calibrate better because they measure facts. The model's own number will become one input among several, not the trigger.

Gatehouse fit

Gatehouse lets you set the confidence level below which a run pauses and goes to a named approver, alongside hard rules that hold a case whatever the score. Cases your approver reviews are kept as labelled examples, which is the raw material for the band table above. Continuous evaluation re-checks against them. The Gatehouse page shows the check in a simulated run.

At a glance

CategoryOperations
Also calledConfidence gating, score threshold, abstention
Borrowed fromStatistical calibration. ML model evaluation
Key sourcesscikit-learn calibration guide. Guo et al. 2017. GPT-4 Technical Report. NIST AI RMF
Typical ownerThe process owner picks the error rate. IT or the vendor measures. Risk reviews drift
The one testOf the cases scored 0.9, how many were actually right?

Sources

  1. [1]Probability calibration, scikit-learn user guidescikit-learn.org In text
  2. [2]Guo et al., On Calibration of Modern Neural Networks, arXiv:1706.04599, 2017arxiv.org In text
  3. [3]OpenAI, GPT-4 Technical Report, arXiv:2303.08774, 2023 (PDF)arxiv.org In text
  4. [4]Kadavath et al., Language Models (Mostly) Know What They Know, arXiv:2207.05221, 2022arxiv.org In text
  5. [5]NIST AI 100-1, AI Risk Management Framework 1.0, January 2023 (PDF)nvlpubs.nist.gov In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io