Volume vs Breadth Signaling How to read a pattern →
6.1 Emerging

Clustering and deduplication for high-volume submissions

Grouping a flood of submissions by what they argue, so a reviewer reads each distinct argument once instead of the same template ten thousand times. It keeps mass campaigns from inflating the record while giving reviewers and the public confidence that no substantive position was missed.

01

The impact of agents

Agencies can't meaningfully review each submission individually once volume reaches hundreds of thousands or millions of written submissions (public comments, but equally consultation responses, grant applications, or planning objections). Mass submission campaigns (historically postcard campaigns, now orchestrated online) produce near-identical submissions that inflate raw counts without adding substantive argumentation. Reviewers need tools that surface distinct arguments rather than re-reading the same template thousands of times.

02

What must be verified

Government needs to read every distinct argument in a body of submissions without re-reading the duplicates that mass campaigns produce. A reviewer can then be confident no unique position was missed when the volume is too high to read each one. Confidence has to rest on what a submission argues, not on judging whether it was machine-written.

03

Protecting access

Semantic grouping tuned to fluent prose can mis-file a submission written in a second language or in plain, unpracticed words. That submission's distinct argument gets folded into a cluster where no reviewer ever reads it. Neither the reviewer nor the submitter can see this happen.

Keeping the path open

  • Deduplicate for analysis and never for exclusion: no submission leaves the record.
  • Publish how the clustering was performed, so participants and the public can verify no substantive argument was lost.
  • Sample clusters by hand, weighted toward the submissions least like the prose the model expects.
04

Response surface

Cluster Console

Submissions are grouped by similarity so a reviewer reads each distinct argument once, with nothing removed from the record.

Foreshore managed retreat plan · submissions
Clustered by natural-language similarity
4,213 submissions · 27 arguments
SUB-2026-0041"My pension hasn’t moved in two years and the acquisition price won’t cover a comparable home…"
SUB-2026-0187template text, signed individually
SUB-2026-0203template text with a personal paragraph added
+ 1,839 more in this cluster
These are unedited samples. Nothing in this cluster was removed from the record.

The console surfaces arguments for a person to weigh. It never scores, ranks for decision, or drops a submission. A three-submission cluster with new data can outweigh a 1,842-submission template; that judgment stays human.

05

Maturity

  1. Emerging Headline

    For deployed comment-analysis tooling, where the CDO Council pilot and Regulations.gov generative-AI processing run in production contexts.

  2. Frontier

    For public transparency of the clustering itself, where participants can verify no substantive argument was lost.

06

Precedents

CDO Council public comment analysis pilot (US). The Federal Chief Data Officers Council, with OIRA and GSA, piloted natural-language-processing tools that recognize topics, group semantically similar submissions, and surface them for subject-matter expert review. The report's worked example collapsed 267 near-identical submissions to 9 distinct comments. The Council went on to publish recommendations for implementing the tools across federal agencies.

Generative AI comment processing on Regulations.gov. ICF, working with GSA, has deployed generative AI to accelerate public comment analysis, moving from spreadsheet-based manual review toward automated clustering and theme extraction. The deployment sits on the US federal rulemaking channel itself.

Delib's Citizen Space analysis tools. The platform offers tagging and coding of qualitative responses, cross-referencing across questions, and AI-assisted first-pass analysis identifying themes and sentiment. It is widely used across UK, Australian, and New Zealand government consultations. The clustering step is already procured infrastructure in three jurisdictions.

The Parliamentary Joint Committee on Human Rights intake taxonomy. The freedom of speech inquiry received approximately 11,460 items and sorted them three ways: 418 accepted as submissions and published, roughly 10,590 categorized as form letters with one sample published alongside the count, and around 452 accepted as correspondence, meaning a view expressed without substantive commentary. The Tasmanian Wilderness inquiry received over 9,600 emails and form letters. A three-way split already separates what is read individually from what is counted from what is sampled.

07

What carries over to agent use

High. Clustering and deduplication are infrastructure-level capabilities that any digital submission system should offer (the comment-analysis precedents transfer directly to grant, objection and petition intake). The CDO Council's approach is designed as a reusable, federal-wide toolset and could be adapted by other jurisdictions.

None of the cited precedents publish their clustering methodology for outside verification. Methodology transparency doesn't transfer from CDO Council or Delib practice; it has to be designed into the deployment fresh.

08

Where things go wrong

Deduplication is analytical, not exclusionary, so on its own it creates no large-scale adverse harm. The safeguard is that clustering surfaces arguments for human review rather than substituting an automated decision for it. The adversarial failure mode is that a language model can paraphrase one campaign into many superficially distinct submissions, defeating deduplication that keys on surface similarity and inflating the count of apparently independent voices. So an agency that treats that count as a tally reads manufactured breadth as genuine support. The durable guard is refusing to treat volume as a vote: weighing submissions by the substance they add so distinct-looking duplicates buy little. Sharper clustering alone doesn't settle this: each detection improvement invites a subtler paraphrase in response, so clustering keeps losing ground on its own.

09

Sources

6 references US · UK · AU