Confidence calibration

When we say 80% confident, is it actually 80%?

AskVerdict AI attaches a confidence score to every verdict. This page tracks whether those numbers hold up against what actually happened, using outcomes people recorded on public debates. No marketing math, just the bucket counts.

Calibration curves from public debates only. Shows predicted confidence vs actual accuracy.

Methodology

How this is measured

What a confidence score is

Every AskVerdict AI verdict includes a stated confidence, for example "80% confident." That's a testable claim, not a vibe. If it's well calibrated, verdicts stated at 80% confidence should turn out correct about 80% of the time, not 50% and not 100%.

How we check it

When a debate has a tracked outcome (marked correct, partially correct, or wrong), we group it into a 10-point confidence bucket, for example 80-90%, and compare the stated confidence against how often that bucket's verdicts actually held up.

Why some buckets say "gathering data"

A bucket with 1 or 2 outcomes can't tell you much, a single wrong call can swing it by 50 points. We only treat a bucket's accuracy as meaningful once it has at least 5 tracked outcomes, and we say so plainly when it doesn't.

What's in scope here

This page shows calibration from public debates only. Outcomes from private and workspace debates are tracked internally to keep improving the underlying models, but they're never mixed into this public figure.

The data

Stated confidence vs. actual accuracy

Each point is one confidence bucket. On the diagonal line means perfectly calibrated: verdicts stated at that confidence turned out correct exactly that often.

We don't have enough resolved outcomes on public debates yet to plot a calibration curve.

This page is wired directly to the tracked-outcomes pipeline, so once public debates accumulate resolved outcomes, the curve above will populate on its own. We'd rather show nothing than show a curve that isn't backed by real numbers yet.

Track your own outcomes

Record what actually happened after a decision and your account builds its own calibration data, separate from the public numbers on this page.