Published benchmarks · Last run

Benchmarks built so you can check every claim.

We publish and maintain benchmarks for the claims we make about community intelligence: what the system reads, how accurately it scores it, and whether the evidence can be checked. Every number traces to a published results file, and the full evidence pack is available for technical review.

Want to check our work? Request the benchmark pack - test data, scoring code, raw outputs, and methodology notes.

Benchmark suites

Defensible reporting v1.0 · Last run

Defensible reporting

Turns consultation evidence into reports reviewers can trace, test, and defend.

See the full benchmark

Report citations resolved

Communiti
372/372
Audit target
100% target

Runs without a decision flip

Communiti
3/3
Audit target
3/3 target

Minority viewpoint preserved

Communiti
3/3
Audit target
3/3 target
Aspect-based sentiment (ABSA) v1.0 · Last run

Issue-level sentiment analysis

Every issue in every response, with its own sentiment, filed under your reporting topics.

See the full benchmark

Issues found in long written submissions

Communiti
100%
Amazon Comprehend
33%

Issues implied but never named

Communiti
90%
Amazon Comprehend
1%

Issues correctly filed under your reporting topics

Communiti
93%
AI assistant
5%
Stance detection v1.0 · Last run

Stance detection

Measures what people want, not just how they sound.

See the full benchmark

Frozen test stance accuracy

Communiti
98.7%
Pass line
85%

Tone-divergent feedback read as stance

Communiti
100%
Tone shortcut
25%

Meaning-preserving wording changes

Communiti
100%
Pass line
95%
Argument mining v1.0 · Last run

Argument mining

Finds the reasons behind community positions, with source evidence attached.

See the full benchmark

Reasons found and correctly filed

Communiti
96.0%
Search shortcut
56.7%

Implied reasons residents never name

Communiti
100%
Search shortcut
37.5%

Quoted arguments rejected by the resident

Communiti
92.3%
Search shortcut
42.1%
Campaign detection v1.0 · Last run

Campaign detection

Find coordinated campaigns without silencing genuine residents who share the same concern.

See the full benchmark

Campaign relationships found

Communiti
96.8%
Duplicate finder
42.9%

Organic pairs correctly kept separate

Communiti
100%
Topic shortcut
81.4%

Unique-voice count accuracy

Communiti
95.7%
Duplicate finder
55.3%

How we benchmark

Numbers you can take to your IT and governance teams

Benchmarks are only useful if you can check them. Ours are built to be checked - and to be rerun when the products change.

Synthetic test data, no resident data

Every corpus is written for testing - spanning typos, sarcasm, voice transcripts, low-literacy writing, implied issues, and ten community languages.

Clear baselines and pass lines

Each benchmark states what was tested, what counted as a pass, and which baseline or comparison helps explain the result.

Reproducible end to end

Every number traces to a published results file, and the full evaluation suite re-runs end to end - including a non-live path that re-scores cached outputs.

Evidence pack on request

Test data, scoring code, raw model outputs, charts, and methodology notes are available for technical review by your IT and AI governance teams.

See your own consultation benchmarked this way

Bring one real, de-identified feedback export to a 30-minute walkthrough and see the evidence trail behind the analysis.

Stay close to the future of community engagement

Product notes, practical field guides, and evidence-led thinking for teams working under public scrutiny.

Read about ourWe care about your data in our privacy policy.

End-to-end engagement workflow