8.4 Frontier

Risk-proportionate review of civic tools

Review depth scaled to consequence: an automated check every tool passes quickly, human review reserved for the tools that can do real harm, and takedown when something slips through. Reserving human review for the tools that can do real harm keeps certification fast enough that developers seek it.

01

The impact of agents

The tools and agents a user might use to deal with government range from a harmless information lookup to something that files a claim or shapes a decision. Reviewing every one to the same depth is neither affordable nor useful, but reviewing none leaves users relying on tools no one has checked.

The gap will only widen as the population of tools gets larger and keeps changing.

02

What must be verified

Government needs a level of review that scales to what a tool can affect: light enough that every tool can pass quickly, deep enough that a tool capable of real harm gets a harder look. The aim is a credible minimum bar, not an exhaustive audit of every tool.

03

Protecting access

The exclusion risk is concentration. An opaque review gate denies the small builder who can't read the rules they're judged against, or can't appeal a rejection. When a queue is built for enterprises, it delays the communities waiting on their tools.

Keeping the path open

  • Publish the review criteria, so anyone can see what is checked.
  • Keep the lightest tier cheap and fast enough for a solo builder.
  • Provide a real appeal route, and never vest review authority in a single entity that can exclude without recourse.
  • Make the submission flow operable by keyboard and screen reader, or a solo builder who relies on assistive technology meets a barrier before the review even starts.
04

Response surface

Submission Review

Review burden scales with consequence, so an informational tool clears automated checks and a decision-support tool goes to a human.

Preview the review pipeline for a different tool category
Your submission · civic registry
  1. Automated checks

    Disclosure label complete · no undeclared data collection · passed in 4 minutes

  2. Human reviewNot required

    Not required for informational tools

  3. Listed

    Live in the registry with its label and mark

Published review criteria

Every check above links to the published criterion it applies. You can read exactly what will be reviewed before you submit, and if your tool is rejected, you’ll be told which criterion it failed.

Review checks behavior against the stated policies. It does not test whether the outputs are true. That is accuracy testing, a separate duty for decision-support tools.

05

Maturity

  1. Established

    As a consumer-app review precedent, documented in Apple's own App Review Guidelines and Google Play's own developer policies: expert review plus malware scanning, with mandatory structured data-collection disclosure.

  2. Frontier Headline

    As a tiered-review model for civic technology, which no jurisdiction has built.

06

Precedents

Apple App Review. Apple's guidelines state that 'every app is reviewed by experts' and that Apple scans 'each app for malware and other software that may impact user safety, security, and privacy', combining automated scanning with human review of app flows, permissions, and usability. App Review evaluated more than 9.1 million submissions in a year and rejected over 2 million. Apple's published review-time metric is that on average 90 percent of submissions are reviewed in under 24 hours.

Google Play's Data safety section. Google relies more heavily on automated machine-learning checks, triggering human review for flagged or sensitive categories, and its developer documentation requires every app to declare how it collects, shares, and protects user data before anyone installs it. The declaration is a condition of distribution, and it is the direct analogue of Apple's disclosure requirement.

Known limitations of App Store and Play review. Malware regularly bypasses review through code obfuscation, delayed execution, and compromised third-party SDKs, with multiple incidents traced to vulnerable advertising SDKs. Store review is a minimum-bar filter and not an exhaustive security audit, and enterprises are advised not to treat presence in a store as sufficient assurance.

07

What carries over to agent use

The app-store model shows that centralized review at scale is feasible but imperfect. Transferable elements: structured disclosure requirements (the Data Safety section is a nutrition label by another name); tiered review (automated first pass, human review for high-risk categories); post-publication monitoring and takedown. Limitations: review checks behavior against stated policies, not ground-truth accuracy; it concentrates review authority in a single entity; and it is binary rather than graded.

A civic registry should adopt the tiered review model but add domain-specific accuracy checks and make criteria publicly auditable.

08

Where things go wrong

App-store review checks behavior against stated policies, not the ground-truth accuracy of outputs. A binary approved/rejected gate of this kind would not catch a method that produces plausible-looking but incorrect results, which pass any behavioral check.

09

Sources

7 references Global · US