T4

Trust Calibration

A user deciding whether to use a government AI agent has little to go on. They may refuse one that would have served them well, or lean on one in a high-stakes case it is not equipped to handle. Calibrating that reliance takes evidence of how far this agent can be relied on for this decision. The agency that fields the agent is the only party positioned to supply it.

Officers inside the agency face the same shortage, with a duty attached that a user does not carry. An officer relying on an agent’s output is making an administrative decision, answerable and on the record. The evidence that lets a user calibrate is the evidence that lets an officer justify that decision, so publishing it serves both.

01

Policy challenge

Increased use of AI agents is happening faster than users have any settled basis for judging when to rely on one. The gap runs both ways: people withhold trust from an agent that would serve them well, and people lean on an agent in high-stakes interactions it cannot be trusted to handle. Automated systems run without calibrated trust or meaningful oversight have already produced serious public harm.

As deployment widens across services, the absence of a shared, legible basis for calibrating reliance becomes the condition users meet by default.

02

Design challenge

Let a user match their reliance on a government agent to the evidence for it.

Provide the means to read what an interaction involves and how much of it is automated.

For complex or high-stakes interactions, build a repeatable, transparent basis for trust.

Let delegation expand on demonstrated use and reverse the moment a user wants it gone.

Keep a path open for people who can't interpret the trust signals themselves.

Patterns in this territory

8 shown
4.1 Frontier

Showing confidence and uncertainty

Showing how sure the agent is, in terms a user can act on: a confidence band tied to what happens next, so a firm determination and a best guess never look alike. A user can tell when an answer needs checking before they act on it.

4.2 Emerging

Trust marks and certification

A trust mark for government AI services that certifies a defined standard the service is held to. The mark names what it stands for, so a user knows at a glance what kind of claim it's making.

4.3 Frontier

Graduated delegation

How much autonomy a user grants an agent over their affairs, growing only as trust is earned rather than handed over all at once. At every point, the user decides how far it goes.

4.4 Emerging

Transparency by default

Telling the user they are dealing with AI, what it can do, and what it is doing right now, at the moments that matter rather than in a terms page. The standing record lets an auditor confirm, after the fact, that the user was told.

4.5 Frontier

Communicating the level of automation

Naming how much of an action a machine decides, in a plain label a user can grasp at a glance. A user reading the label knows whether a person or a machine settled the outcome they're about to rely on.

4.6 Emerging

Signaling human-in-the-loop oversight

Making it clear, at every point in an agent-run decision, that a person reviews it. A user can tell whether that review is real.

4.7 Frontier

Earned trust and reversible delegation

Delegation that expands on demonstrated use and contracts the moment the user wants it gone. A user can withdraw delegation at any point without losing access to the outcome they came for.

4.8 Frontier

Declared automation on outbound decisions

The outbound decision itself declares what was automated and to what degree, in a form both the person and their agent can read. Reliance can then be calibrated decision by decision.

Case studies that touch this territory