Consensus you can audit
Ask a research question. A panel of AI personas across several providers each takes a position, every claim is checked against retrieved literature, the panel rates each other — and you get the agreed answer with the dissent still attached.
The panel is not open to public questions yet. The engine, its claim verification and its epistemic label are built and audited — this surface opens when the panel is switched on.
How it works
Draft. Verify. Rate. Publish.
A panel takes positions
Independent personas across several model providers each answer the question on their own, before seeing anyone else's answer. No single provider can carry a panel — seats are capped per provider.
≤2 seats per providerClaims are checked
Each position is decomposed into atomic factual claims, and each claim is checked against literature retrieved at answer time. Claims that fail retrieval are marked unverified rather than quietly dropped.
retrieval-backed verificationThe panel rates itself
Personas rate each other's positions in a structured round. A persona cannot rate its own provider's answer, so agreement is measured across providers, not within one.
same-provider self-rating bannedConsensus — and dissent
The agreed position is synthesized and then re-verified as a statement in its own right. A minority position that survives the rounds is published alongside it, never averaged away.
dissent is preservedThe epistemic label
Every answer says how much to trust it.
An answer without its uncertainty is a liability. Each published consensus carries the measurements behind it — and where a measurement was not taken, it says so instead of showing a zero.
- Panel claims verified
- How many of the panel's factual claims were confirmed against retrieved sources. A low ratio is shown as a low ratio.
- Statement re-verified
- A set of true claims can be summarized into a false one, so the published statement is decomposed and checked again on its own.
- Panel agreement (W)
- Kendall's W over the panel's ratings. Shown as “not computed” when too few raters make it meaningless.
- Cross-provider agreement
- Agreement measured only between different providers — the number that cannot be inflated by one vendor agreeing with itself.
- Evidence overlap
- How much the panelists' evidence sets actually overlapped. Low overlap means they reasoned from different sources.
Limits
What this is not.
- This is statistical agreement among AI models. It is not human expert consensus, peer review, or professional advice.
- Verification depends on what public literature retrieval returns at answer time. An unverified claim is not necessarily false — it means retrieval did not support it.
- Models share training data and can be wrong together. Cross-provider agreement reduces that risk; it does not eliminate it.
- Questions asking what a specific person should do are declined. So are purely value-laden questions, which evidence cannot settle.
- Every statistic on an answer is measured or shown as absent. Nothing is filled in with a default.
Provenance
Delphi Commons is the free, public surface of OpenDelphi, a consensus-driven research instrument platform. The same Delphi machinery that converges expert panels on what to measure runs here in the open, on questions anyone can ask.