Research notes

Findings from the data, and how we measure.

Everything here is built from consented, aggregated results — never from a text anyone pasted in, and never about one learner, however anonymised. Cohorts are published at 200 consented results or more. Where we do not have the data yet, the entry says so.

The method and its 32 sources · Quarterly data releases · [email protected]

Index

Research notes · vigilance

Most people miss the mistake: what 10,000 benchmarks show

The finding in one sentence: across 10,412 consented benchmarks, readers marked the planted error in 38% of texts — and readers who found it were no slower than readers who did not.

If you read AI-written text at work and assume you would notice a wrong number, this is your baseline.

Catch rate by error type · N = 10,412 · one planted error per text

Each bar is the share of readers who marked the planted sentence when that error type was present. False alarms are counted separately and are not subtracted here.

The pattern is consistent and, we think, explainable. A wrong number sits in one place and contradicts something nearby, so a careful reader can find it by comparing two lines. A wrong conclusion does not contradict any single line — it contradicts the reasoning above it, which means holding several sentences in mind at once while reading a text that is already dense.

Omitted conditions were caught least often of all. Nothing in the text looks wrong; something that should be there simply is not. Readers cannot compare against a sentence that was never written.

Reading speed did not predict catching the error. The readers in the fastest quarter caught it at 37%, the slowest quarter at 39% — a difference inside the noise. Care and speed came apart, which is the reason we report them as two numbers and never blend them into one score.

What it doesn't mean

This is a benchmark, not a workplace. Readers knew they were being measured and still missed the error most of the time, so the workplace number is probably worse — but we cannot say by how much, and we will not guess.

It also does not show that error-catching can be trained. That is the open question, and it is in the "we're testing" column on the method page. We have a baseline; we do not yet have a before-and-after.

One error per text is an artificial rate. Real AI output sometimes has none and sometimes has several.

Instrument
Benchmark v1.4: one 280–500-word AI-generated passage at a fixed reference difficulty, ten four-option questions (seven literal, three inferential), one planted error detectable from the passage alone.
Readers
10,412 consented benchmarks, 8 first languages, 6 fields, B1 and above by self-report.
Period
March – August 2026.
Consent basis
Opt-in at the result screen, unticked by default, revocable. Screen-reader sessions are excluded from norms.
Limits
Self-reported level; one error type per text; benchmark conditions, not work conditions; no control group.
Version
Published 11 September 2026. No corrections to date.

Three minutes, free: take the benchmark.

Quarterly data releases

Aggregates under an open licence with attribution. Each release lists the cohorts it covers, the instrument version and the consent basis. Cite this page.

ReleaseCohortsInstrumentResultsFile

Citation: dub.mom research notes, quarterly aggregate release, <quarter>. Retrieved <date>.

Requesting other aggregates

We will prepare an aggregate that is not in a release if the cohort is 200 consented results or more and you name the purpose. Write to [email protected].

Never released

  • Anything derived from a text a learner pasted in — those are not stored and not analysed
  • Anything about a single learner, however anonymised
  • Any cohort smaller than 200 consented results
  • Percentiles for an individual, or a name next to a number
  • Team-level data identifying a company without that company's written agreement