Showing confidence and uncertainty
Showing how sure the agent is, in terms a user can act on: a confidence band tied to what happens next, so a firm determination and a best guess never look alike. A user can tell when an answer needs checking before they act on it.
The impact of agents
Agents that produce determinations of varying certainty are now widely available, and as government services run on them they will issue far more of these determinations than any caseworker did. Presenting them all with uniform authority drives either uncritical acceptance (automation bias) or blanket rejection (automation aversion).
When the certainty behind an output is invisible, a user cannot tell a confident determination from a guess, so they cannot calibrate how far to rely on it.
What must be verified
Government needs the certainty behind a determination to be legible and accurate before anyone relies on it, so reliance can be calibrated rather than driven by over-trust or blanket rejection. The agency issuing the determination must hold and expose that confidence calibration alongside it.
Protecting access
A user who can't see a color band, or doesn't read percentages and traffic-light conventions, gets none of the meaning the indicator conveys. They may over-rely on a guess or reject a sound determination that a better-served user would have calibrated correctly.
The routing behind the band excludes differently. A person whose circumstances are complex or atypical, several income sources, a thin record, a name the system matches inconsistently, scores low confidence more often and is redirected to the human-review queue. When that queue runs slower than the automated path, the people hardest to score wait longest for the same outcome.
Keeping the path open
- Pair every confidence signal with a plain-language equivalent ('we're fairly sure about this, but a person will double-check') and a next step the user can act on.
- Hold the human-review path to the same timeliness standard as the automated one, with both processing times published.
- Convey the same meaning non-visually: every badge needs an accessible name, and the band must stay legible at high zoom.
Response surface
The system's own certainty about a determination is shown to the user in bands, each one paired with what happens next at that level.
Housing assistance: you appear eligible to apply
An officer will review this before any decision takes effect.
Low confidence. 0 of 3 confirmed. An officer will review this before any decision takes effect.
Maturity
- Established
For confidence surfacing in clinical decision support and weather forecasting, where the response and the theory of appropriate reliance are mature.
- Emerging
For government digital services generally, where no specific guidance requiring confidence display is cited here; the healthcare and forecasting patterns above are the nearest documented analog.
- Frontier Headline
As applied to a user-facing government determination, where the response remains unproven.
Precedents
Lee and See on trust in automation. The framework defines calibration as the correspondence between a person's trust in an automated system and that system's actual capabilities, and names miscalibration in both directions: overtrust produces misuse, distrust produces disuse. It asks for two further properties, resolution, or how precisely a judgment of trust differentiates levels of capability, and temporal and functional specificity, so that trust attaches to a particular function at a particular moment. Calibration alone is not the requirement.
Clinician trust in AI diagnostics. Clinical systems have developed confidence-visualization patterns: color-coded bands, uncertainty intervals alongside predictions, and low-certainty labels that trigger escalation. Clinicians are generally receptive to evidence-based AI tools, and override rates stay high where calibration is poor.
Amershi et al., Guidelines for Human-AI Interaction. The guideline directs a system to make clear how well it can do what it can do, from a set validated with practitioners against real interfaces. The capability claim is treated as part of the interface, and not as documentation sitting beside it.
What carries over to agent use
High transferability. Government services regularly produce determinations of varying certainty (eligibility assessments, risk classifications, benefit calculations), so surfacing confidence is directly applicable. The healthcare parallel is apt: clinicians and caseworkers both need to know when to rely on a system versus apply professional judgment. The color-coded band pattern is simple and well understood.
Key adaptation: government confidence signals have to be tied to a next step the user can act on ("This assessment has medium confidence; a human officer will review before any decision takes effect"), not displayed as passive information the user can do nothing with.
Where things go wrong
The failure mode is an automated estimate issued with false certainty, masking determinations that should never have stood. Surfacing low confidence on such an estimate, tied to mandatory human review, flags exactly those determinations before they go out at scale. The confidence score itself can be gamed by the agency that relies on it: a threshold loosened to shrink the review queue reports higher certainty than the estimate supports, and the display never reveals the difference.
Sources
3 references
The instrument, the operating deployment, or the official record itself.
- Lee, J.D. & See, K.A. — Trust in automation: Designing for appropriate reliance
-
Amershi et al. — Guidelines for Human-AI Interaction (CHI 2019)
Guideline G2, 'Make clear how well the system can do what it can do,' from an 18-guideline set validated by 49 practitioners against 20 AI products, a peer-reviewed anchor for surfacing error rates and confidence. Validated against 2018-era classification and recommendation interfaces, not autonomous agents, so it transfers as a principle and carries no evidence tested on agents.
Writing about the subject rather than the framework itself, including vendor commentary.