Measuring the Human Cost of Verifying AI Output
In one line
When you could do a task yourself, does checking an AI's attempt cost less time than doing it? And how often does a wrong answer get through?
Links
- Live site: https://verification-cost.pages.dev
- Proposal as one page: https://verification-cost.pages.dev/proposal.html
- Proposal as PDF: https://verification-cost.pages.dev/proposal.pdf
Status
Proposal for the Lossfunk Fellowship, October 2026. Not yet run: the study needs ethics approval and a panel first.
Why I can run this
- I learn what I need and ship it. Before taking any ML courses, I taught myself enough to build an unsupervised anomaly-detection pipeline for ship-tracking data end to end, and published it as first author (DoSCI-2026).
- I build the measurement, not just the system. For my Contract Risk Analyzer, which I designed, built and deployed myself, I wrote my own 90-clause benchmark to test whether it actually worked.
- I scope to what's real. This proposal was rebuilt twice: once after a literature review narrowed the claim, and once to fit the 3–4 people I can actually count on. Every figure from prior work was checked against the original papers.
The problem
In one sentence
When a person could do a task themselves, does checking an AI's attempt cost less of their time than doing it, and how often does a wrong AI answer get through the check?
The problem
Delegating work to AI assumes that checking its output is cheaper than doing the work. If checking takes about as long as doing, or wrong answers slip through, delegation saves little or nothing.
Other fields have measured this. Machine-translation research has timed the same translators translating from scratch versus correcting machine output (Plitt & Masselot 2010; Green et al. 2013; Sarti et al. 2022), and a 2024 radiology pilot did the same with deliberately planted AI errors (Acosta et al. 2024). But in those studies people rewrite the output, and quality is judged by humans, not against an exact answer.
In code and data analysis, the studies I found measure only part of it: - Checking alone. Programmers judging LLM-written assertions accepted about half of the wrong ones, and decided on wrong accepts faster than on correct rejects (Kaufman et al. 2026). - Doing vs. AI-assisted work, in different groups. SQL users with an AI tool finished about 31% faster than users writing by hand, but produced roughly the same number of correct answers per hour, and were somewhat less accurate (46% vs. 64% of queries correct, not statistically significant) (Ipeirotis & Zheng 2025). - With vs. without AI, end to end. Developers were 19% slower with AI while believing they were faster (METR 2025); AI code reviews saved no time (Tufano et al. 2025).
Crescitelli et al. (2026) proposed a protocol for measuring verification cost against a time budget, but ran no experiment.
The gap: in this search I did not find a study in code or data analysis that times the same people solving tasks and judging AI attempts at matched tasks, with automatically checkable answers, and reports check time, do time, and the rates of wrong answers accepted and correct answers rejected.
Key definitions
- Do time: minutes for a person to produce an answer themselves. Correctness is recorded against the answer key.
- Check time: minutes from seeing the AI's answer to giving a verdict, within a fixed budget.
- Verdict: "correct" or "incorrect", plus confidence (50%, 75% or 100%).
- Outcomes of a check: Crescitelli et al.'s four outcomes: correct verdict; wrong verdict; budget ran out; declared unverifiable. Wrong verdicts are split into false accepts (wrong answer judged correct) and false rejects (correct answer judged wrong).
- Time to correct verdict: check time counted only when the verdict was right. A fast wrong verdict is an error, not a fast check.
- Verification ratio: check time ÷ do time, on matched tasks. Both are human-minutes, so the ratio is meaningful. The primary estimate comes from a survival model, so timed-out attempts count as "at least this long"; raw medians of finished attempts are a secondary estimate.
- Premature closure: a 100%-confidence verdict that is wrong.
- Delegation break-even: delegating saves time only if
check time + (non-accept rate × do time) < do time. A non-accept is any check that doesn't end in trusting the answer: a reject (whether the answer was right or wrong), the budget running out, or "unverifiable". Each leaves the person redoing the task by hand. False accepts are reported separately as a quality cost, because their price depends on where the output is used.
Scope and setting
- Core study: code only. Data analysis is a later phase.
- Code: short Python functions (about 15–40 lines) with a written specification and a hidden test suite. Tasks are newly written, to avoid models having seen them in training.
- Data analysis (later phase): a small CSV table (about 50–200 rows) and a question with one exact numeric answer, computed in advance by script. The AI's answer includes its pandas code and the final number, so checkers can read and run its steps rather than only recompute the number.
- Task size: about 5–10 minutes to do by hand. In the core study, 5–10 minutes is a writing target, checked against observed do times. In the full study, pilot 1 (timing) sets it.
- Checkers: a panel of 3–4 people from the Lossfunk community with working Python; a short screening task filters out people below a minimum level. Small-N: each person is treated as a replication, and claims are about this panel only.
- Allowed while checking: reading, running the code, writing their own tests, a calculator. Not allowed: AI assistants or web search.
- What checkers do: give a verdict only. They do not fix the answer. This separates the cost of judging from the cost of rewriting, which translation and radiology studies combine.
- Mix: half of all answers are wrong; checkers aren't told the ratio. In the core study, each session has at least 1 correct and 1 wrong answer; in the full study, at least 2 of each.
- Error sources: in the core study, planted and natural errors both serve as wrong answers for RQ2; natural errors are also described per model for RQ3.
- Planted errors, three types per domain. Code: boundary or off-by-one, wrong condition, missed edge case (empty input, duplicates). Data analysis: wrong filter or dropped rows, wrong aggregation or grouping, rounding or unit. Each at two levels: obvious (fails a simple case, or gives an implausible number) and subtle (passes typical cases, or is close to the right number).
- The models' own errors, sampled from 2–3 models with pinned versions and settings.
Research questions
- Time. On matched tasks, how does check time compare with do time, and how does the verification ratio change with difficulty? Where does delegation fail to break even?
- Errors that get through. How often are wrong AI answers accepted within budget, and how does that depend on subtlety? In the core study this is descriptive: acceptance by subtlety and by error type, with counts.
- Across models (exploratory). For errors the models make on their own, are a stronger model's wrong answers caught less often, or more slowly, than a weaker model's?
- AI checkers (supporting). Do AI checkers miss the same errors that human checkers miss?
Hypotheses
- H1 (per person): For each person, checking is faster than doing on average (ratio below 1), but the ratio rises with task difficulty. On the hardest code tasks, delegation will fail to break even for a meaningful share of checkers. In the core study, H1 is tested descriptively against difficulty labels (easy, medium, hard) set when the tasks are written and checked against observed do times. Why: checking non-trivial code often means re-deriving its logic; METR and Ipeirotis & Zheng show time savings do not automatically become correct output.
- H2: Subtle errors are accepted much more often than obvious ones, especially missed edge cases. Why: Wen et al. and Kaufman et al. show plausible wrong answers pass time-limited checkers. On experience: With 3–4 people, experience effects can't be tested.
- H3: Wrong answers are accepted faster than correct answers are rejected. Why: checkers who accept stop looking, while checkers who reject must find the problem. This replicates Kaufman et al.'s comparison in a new setting, and extends it by comparing both against each person's own do time.
- H4 (exploratory, uncertain): A stronger model makes fewer errors, but each one is more likely to get through. Its total checking burden may still be lower, because there are fewer errors to catch. Why uncertain: Kirchner et al. show checkability depends on how a model was trained, not only on capability.
- RQ4 (open question, no prediction): Do AI checkers miss the same errors that human checkers miss?
What would change my mind
- The question is too easy here. If the ratio is well below 1 on nearly all tasks and false accepts are rare, these tasks don't stress checking. I would move to harder or longer tasks, or to studying what makes checking faster.
- Recruitment fails (full study). The core study uses "Fewer than 3 people" below instead. If the pilots show I can't get 3 verdicts per AI answer, plus 6 for a small exploratory subset of wrong answers, I would first drop the extra RQ3 answers and the exploratory subset, then use fewer tasks per domain, and only then drop to one domain. On the subset, the protocol can only classify an error at six verdicts when every checker misses it or none do; anything in between is inconclusive (Crescitelli et al. 2026, Section 5).
- Natural errors are too rare. If the stronger models make too few errors on these tasks, I report RQ3 as a limitation. I would not swap in planted errors and still call it a model comparison.
- Planted errors feel fake. In the full study, if checkers in pilot 2 (dress rehearsal) rate planted errors at least 1 point lower on average (1–5 scale) for being AI-written than the models' own errors, I would rebuild them from real model errors. In the core study nothing can be rebuilt afterwards: if the realism ratings show the same gap, I report planted errors separately and base conclusions mainly on natural errors.
- Fewer than 3 people. If fewer than 3 people complete all their sessions, the human part becomes a pilot, and the instrument, task set and AI-checker comparison become the main results.
Out of scope
- Data analysis in the core study. It moves to a later phase.
- Effects of experience. They can't be tested with 3–4 people.
- Claims about programmers in general. Results describe the panel.
- Fixing or rewriting AI output. Checkers only give verdicts; that's the difference from post-editing studies.
- Long, multi-step agent work (hours, several agents at once). METR found timing unreliable when people run agents in parallel.
- Domains without automatic ground truth, such as open-ended writing or legal review.
- Interactive checking, where the checker questions the model. It adds persuasion effects (Randazzo et al. 2025) and is a separate study.
- Making models easier to check. That is Kirchner et al.'s agenda. I measure checkability; I don't change it.
- Causal claims about capability. RQ3 compares a few models and cannot show that capability itself causes higher verification cost.
References
- Acosta et al. 2024. The Impact of AI Assistance on Radiology Reporting: A Pilot Study Using Simulated AI Draft Reports. arXiv 2412.12042
- Becker et al. (METR) 2025. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv 2507.09089
- Crescitelli et al. 2026. AI Evaluation Should Measure Verification Cost, Not Correctness Alone. arXiv 2608.08709
- Green, Heer & Manning 2013. The Efficacy of Human Post-Editing for Language Translation. CHI 2013. DOI 10.1145/2470654.2470718
- Ipeirotis & Zheng 2025. Natural Language Interfaces for Databases: What Changes for SQL-Literate Users? arXiv 2511.14718 (v3, 2026)
- Kaufman et al. 2026. Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions. arXiv 2607.08885
- Kirchner et al. 2024. Prover-Verifier Games Improve Legibility of LLM Outputs. arXiv 2407.13692
- Plitt & Masselot 2010. A Productivity Test of Statistical Machine Translation Post-Editing in a Typical Localisation Context. PBML 93. https://ufal.mff.cuni.cz/pbml/93/art-plitt-masselot.pdf
- Randazzo et al. 2025. GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs. Harvard Business School Working Paper 26-021. https://www.hbs.edu/ris/Publication%20Files/26-021_8db29bd1-04ed-4e98-86de-d286721afc7a.pdf
- Sarti et al. 2022. DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages. EMNLP 2022. https://aclanthology.org/2022.emnlp-main.532/, arXiv 2205.12215
- Tufano et al. 2025. Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword? ICSE 2025. arXiv 2411.11401
- Wen et al. 2024. Language Models Learn to Mislead Humans via RLHF. arXiv 2409.12822
What's known
Papers
Doing vs. correcting AI output, in other fields
The same people do a task from scratch and correct machine output, both timed.
| Paper | What they did | What they left open |
|---|---|---|
| Plitt & Masselot 2010 (PBML 93) PDF | 12 professional translators from Autodesk's localisation vendors, timed translating from scratch vs. post-editing machine translation. Post-editing saved about 43% of translation time. | People rewrite rather than judge. Quality rated by humans, not an exact key. No measure of errors let through. |
| Green, Heer & Manning 2013 (CHI, DOI) | 48 professional translators, 3 languages. The same people translated and post-edited, on different sentences. Post-editing was faster and produced better quality. | Same as above. |
| Sarti et al. 2022 (DivEMT, EMNLP) ACL Anthology | 18 professional translators, 6 languages, keystroke and timing logs. Speed-ups from post-editing ranged from about 10% to about double, depending on language and translation system. | Same as above. Shows the benefit depends heavily on output quality. |
| Acosta et al. 2024 (arXiv 2412.12042) | 3 radiology readers (one radiologist, two residents) wrote reports from a template or edited AI drafts, with errors planted in half the drafts. Median reporting time fell from 573 to 435 seconds. | Tiny sample. Expert-judged, not an exact key. Editing, not a verdict. |
Checking AI output only
People judge AI output, but never do the task themselves.
| Paper | What they did | What they left open |
|---|---|---|
| Kaufman et al. 2026 (arXiv 2607.08885) | 86 Python programmers judged LLM-written assertions. They accepted about half of the wrong ones, and decided on wrong accepts faster than on correct rejects. | No do-it-yourself comparison. Judged small assertions, not full solutions. |
| Wen et al. 2024 (arXiv 2409.12822) | Time-limited checkers (3–10 min) judged answers from one model before and after RLHF. Accepting wrong answers rose by 24.1 points on reading questions and 18.3 on code. | No do-it-yourself comparison. One model, before vs. after training. |
| Kirchner et al. 2024 (arXiv 2407.13692) | Checkers had 45 seconds per maths solution. Training for correctness only made solutions slower and harder to check; training for checkability reversed that. | No do-it-yourself comparison. Easy school maths, very short budget. |
| Mozannar et al. 2024 (CHI, arXiv 2210.14306) | 21 programmers using Copilot labelled their own screen recordings. Verifying suggestions took 22.4% of their time. | Autocomplete only. No record of wrong code accepted. |
With vs. without AI, end to end
Total time is compared, but checking isn't separated out.
| Paper | What they did | What they left open |
|---|---|---|
| Becker et al. 2025 (METR, arXiv 2507.09089) | 16 experienced developers, 246 real tasks randomly allowed or disallowed AI. They were 19% slower with AI while believing they were faster. About 9% of AI time went to reviewing and cleaning up AI output. | Total time only. No separate checking step. No record of wrong output accepted. |
| Tufano et al. 2025 (ICSE, arXiv 2411.11401) | 29 professional developers reviewed code manually, with a ChatGPT review, or with a secretly "perfect" review. Neither AI condition saved time; manual reviews were, if anything, faster, though the differences weren't significant; AI reviews made people focus only where the AI pointed. Reviewers found a median of about half the injected issues, with or without the ChatGPT review. | Review comments, not verdicts on solutions. One model, small sample. |
| Ipeirotis & Zheng 2025 (arXiv 2511.14718, v3 2026) | 20 SQL users on tasks with known answers: half wrote by hand, half used an AI tool. The AI group was about 31% faster, but produced roughly the same number of correct answers per hour, and were somewhat less accurate (46% vs. 64%, not statistically significant). | Different people in each group. No deliberately wrong queries. |
Protocol and theory
| Paper | What they did | What they left open |
|---|---|---|
| Crescitelli et al. 2026 (arXiv 2608.08709) | Defined Verification-Cost Errors (wrong answers enough checkers miss within budget) and a six-step protocol to measure them. | Never run. No do-it-yourself comparison. |
| Vasconcelos et al. 2023 (CSCW, arXiv 2212.06823) | 731 people across 5 studies. People over-rely on AI more when checking it is costly. | The theory behind the question, not a do-vs-check measurement. |
The gap
Other fields have timed the same people doing a task versus correcting machine output, and code studies have measured checking alone or doing versus AI-assisted work in separate groups. In this search I did not find a study in code or data analysis that times the same people solving tasks and judging AI attempts at matched tasks, with automatically checkable answers, and reports wrong answers accepted and correct answers rejected.
What I borrow
- Same people, both conditions, different items: Green et al. 2013.
- Planted errors with some outputs left correct: Acosta et al. 2024; Tufano et al. 2025.
- Time limit, gold answers, confidence rating, checkers blind to which model produced an answer: Wen et al. 2024.
- Time to a correct verdict: Kirchner et al. 2024.
- Wrong accepts faster than correct rejects (to test, H3): Kaufman et al. 2026.
- Correct answers per hour, not just speed: Ipeirotis & Zheng 2025.
- Budget, four outcomes per check, premature closure, several checkers per answer: Crescitelli et al. 2026.
- One task at a time, no parallel agents (timing breaks otherwise): METR's February 2026 design update.
The study
Core study (what I'll run)
Overview
A panel of 3–4 people from the Lossfunk community each solve 12 code tasks themselves and judge AI answers to 12 matched tasks. AI models also judge every AI answer, as a supporting comparison.
flowchart TD A["Write 32 code tasks (24 used, 6 spare, 2 practice)"] --> B[Model answers: 1 correct and 1 wrong per task] B --> C[Pair tasks] C --> D[Do: solve one task of each pair] C --> E[Check: judge the AI answer to its partner] P["Panel: 3–4 people, sessions of 3 pairs (max 60 min)"] --> D P --> E D --> G[Do time and correctness] E --> H[Check time, verdict and confidence] G --> I[Per-person ratios and false accepts] H --> I B --> J["AI checkers: 3–5 models judge every answer 5 times"] H --> K[Errors missed by humans vs. AI checkers] J --> K
Domain
- Code only. Data analysis moves to the full study / a later phase.
Panel
- 3–4 people from the Lossfunk community with working Python, after the same screening test. I (the author) don't take part, since I wrote the tasks and know the errors.
- This is an assumption to confirm with Lossfunk. If more people become available (e.g. a few from each residency cohort), the same design scales.
- Small-N design: each person does many tasks, and each person is treated as a replication. Results describe this panel, not programmers in general. Crescitelli et al. say the same: where verifiers aren't sampled representatively, estimates are "statements about the observed verifier panel" (Section 6).
- Screening: screen about 5 people to get 3–4. If fewer than 3 pass, the "Fewer than 3 people" rule in PROBLEM.md applies.
- Reserve rule: if someone drops out before their first session, a reserve from screening takes their place in the allocation. After their first session, no replacements.
Tasks and time
- About 3 hours per person: 4 sessions of about 45 minutes, plus a short practice session on 2 practice tasks, written separately and not used in the study.
- Each person sees 24 tasks, each exactly once: 12 they do, 12 they check.
- Write 32 tasks: 24 used, 6 spares (for dropped tasks), 2 practice.
- Task writing: I draft every task. AI may only review specs and suggest edge-case tests.
- Disagreements: if the second person's solution disagrees with the key, the spec or tests are fixed and re-checked. If it can't be resolved, the task is dropped and counts against the spares.
- Difficulty: when writing tasks, I label each easy, medium or hard, with 8 of each among the 24 used. The labels are checked against observed do times.
- Pairing: tasks are paired within the same difficulty label, matched by my estimated do time within each label: 4 pairs per label, 12 pairs in total.
- If people can give about 5 hours, expanding to 40 tasks per person is a separately pre-registered follow-up after the core sessions: write 16 new tasks and use them, so the 6 original spares stay available for the realism check. The 16 are labelled so the 40 used tasks stay roughly balanced across easy, medium and hard. Each label needs an even count so pairs stay within a label (e.g. 6, 6 and 4 new, giving 14, 14 and 12).
- Each session: 3 pairs (do one task, check the AI's answer to its partner). Order randomised; do-first and check-first alternate across sessions.
- Each session ends after 3 pairs or 60 minutes, whichever comes first. Pairs not started move to the next session, adding a fifth session if needed.
AI answers
- 2 answers per task: 1 correct, 1 wrong. Each checker sees one.
- The wrong answer is a natural model error if one exists, otherwise a planted error. Aim for about half natural and half planted overall.
- 50% of checks are wrong answers; each session has at least 1 correct and 1 wrong.
- Models: one strong closed model, plus one or two open-weights models of different sizes; exact versions and settings recorded.
- Model answers: each model answers each task 5 times, offline.
Where code runs
- Participants run their own code and the worked examples in the browser. Hidden tests are never sent to the browser.
- Submitted answers are graded after the session, offline on my machine, in the same sandbox used for model answers: a container with no internet and a time limit.
- Participants get no right/wrong feedback during the study.
Allocation
- Each task is done by about half the panel and checked by the other half, with correct and wrong answers split between checkers. Use a balanced Latin-square scheme and report the actual allocation.
- With 4 people: each task done by 2, checked by 2 (one sees the correct answer, one the wrong). With 3 people: 1–2 and 1–2, balanced across tasks.
Timing (no pilot participants available)
- One fixed check budget for all tasks: 10 minutes, the top of the 5–10 minute target do time, so a check that uses all of it saves nothing on a typical task. This differs from the full study, where the budget is each task's median pilot 1 do time. Do-time cap: 20 minutes. Tasks are written to take 5–10 minutes to do.
- An informal walkthrough of the harness with the author and one other person tests the software only. It is labelled as such and stores no timing or verdict data.
Realism
- Before the study, in a separate, consented step, the walkthrough person rates a sample of answers blind: "How likely is it that an AI wrote this exact answer?" (1–5). These ratings are stored and used.
- After each person's final session, they rate about 8 answers to spare tasks, which nobody in the panel has seen, with the same question. If fewer than 4 spares remain after drops, they rate answers to tasks they did, and this is noted. This comes last so it can't affect how they check.
- Nothing can be rebuilt after the study, so the result is used to interpret: if planted errors are rated at least 1 point lower on average than natural ones, they are reported separately and conclusions rest mainly on natural errors.
- Held-out set: the spares used for this check are the held-out set. Their task text stays private; their realism ratings are released with task IDs and error types, but not the task text. If drops push the realism check onto tasks people did, the held-out set is whatever spares remain, possibly none; the release says so.
- This check is indicative, not conclusive: about 8–10 ratings per kind of answer, plus the walkthrough person's ratings. The spare-task ratings come from roughly 3 distinct answers per kind, repeated across raters.
Analysis
- Per person first: each person's verification ratio, with intervals from resampling their tasks. Report how many of the 3–4 people show each pattern.
- H1, descriptively: each person's ratio for easy, medium and hard tasks, using the difficulty labels.
- Pooled: person as a fixed effect, task as a random effect. Survival-style model for timeouts, as before.
- RQ2 (descriptive): false-accept rate with an interval; subtle vs. obvious and error type descriptive, with counts.
- H3 (wrong accepts faster than correct rejects) per person.
- RQ3 (descriptive): per model, its natural errors and how often humans caught them. Expected about 4–6 natural errors per model.
- Session trend: session number and do-first/check-first order, reported as a session-by-session trend.
- Guessing: nobody is excluded. Verdicts given very fast (under 30 seconds) at 50% confidence are flagged as possible guesses, and results are reported with and without them.
- No VCE classification: it needs at least 6 checkers per answer. With 4 people each answer gets 1 human verdict; with 3 people, about half the tasks have only one of their two answers checked. Report per answer whether its checker caught it. Error rates are estimated by pooling across answers.
- With 3 people × 12 checks, about 18 wrong answers are checked in total, so results are pilot-sized.
AI checkers (supporting part)
- 3–5 AI models each judge every AI answer (correct and wrong) except answers they wrote themselves. They get the same spec, worked examples and instructions as humans, but no code execution (text-only judgement): verdict plus confidence.
- Each model judges each answer 5 times, to measure consistency.
- Compare with humans on the same answers: which errors both miss, which only humans miss, which only AI checkers miss.
- This does not measure verification cost (no human time). It asks whether an AI checking first could reduce human checking cost.
- A version with code execution is a possible extension.
Pre-registration
- The core study is pre-registered on OSF before the panel. This includes one contingency: if sessions often run past 60 minutes, move to 2 pairs per session and add sessions; the pairs and allocation don't change.
- AI checkers run after pre-registration.
Ethics, payment and consent
- Approval before any data is collected from people.
- Panel members volunteer; no fee. They can stop at any time. With their agreement, they're acknowledged in the write-up. If Lossfunk offers a small budget, it goes to vouchers to widen the panel.
- This overrides the full study's payment and consent lines.
Full study (if resourced)
The design for a larger study, kept for when more participants are available. Where it differs, the core study above takes precedence.
Overview
Participants solve some tasks themselves and judge AI answers to matched tasks. Everything runs in a web harness that shows one task at a time and records timing, code runs and verdicts.
Order of steps
- Ethics approval.
- Write 26 tasks per domain, with specs written against the edge-case checklist (see Tasks and domains); a second person checks each.
- Generate model answers.
- Drop tasks no model solves; keep 24 per domain. If 0 or 1 tasks are dropped, the 24 kept are chosen at random, with the random seed recorded (do times aren't known until pilot 1, so they can't be used).
- Plant errors.
- Pilot 1 (timing): set budgets and difficulty.
- Pair tasks.
- Pilot 2 (dress rehearsal): test harness, allocation, session length; realism ratings.
- Finalise allocation.
- Pre-register on OSF.
- Main study.
- Analysis.
Pairing comes after dropping, so a dropped task never breaks a pair.
Tasks and domains
- Code: a Python function of about 15–40 lines, with a written specification and two worked examples. A hidden test suite of at least 20 tests, including edge cases, decides correctness.
- Data analysis: a CSV table of about 50–200 rows and a question with one exact numeric answer, computed in advance by script. The question states the required rounding, and answers are marked against it.
- New tasks only. I write every task; a second person solves each one to check the specification is clear and the answer key is right. This avoids tasks models may have seen in training.
- Edge cases in the spec. Specs are written against a fixed checklist: empty input, duplicates, ties, negative values, boundary values, and missing data (for data tasks). The second-person check confirms the spec states the behaviour for each item that applies. Planted errors may only target edge cases the spec already states.
- Task size: about 5–10 minutes to do. This is an initial target; pilot 1 (timing) sets the final value.
- Pilot 1 (timing): about 18 people, do only, aiming for 3 do attempts per task (48 tasks × 3 = 144 attempts, about 8 tasks each). It sets difficulty, pairs and check budgets. Medians from 3 people are noisy.
- Pilot 2 (dress rehearsal): about 8–10 people, after budgets are set. Full timed sessions with the real procedure, harness and allocation script, except that natural errors are oversampled so each kind of answer gets about 15 realism ratings. It tests the harness, the allocation, session length and budgets, and collects realism ratings (see Planted errors). Its data isn't used in the main analysis.
- Pilot recruitment: about 26–28 people in total. Participants from both pilots are excluded from the main study, and pilot verdicts don't count toward the main study's 42 sessions per domain.
- Difficulty. Pilot 1 (timing) measures each task's median do time. Tasks are grouped into easy, medium and hard by that time.
- Pairs. Tasks are paired within a domain by pilot 1 do time (within about 20% of each other) and by the kind of skill they test.
- Pool size (initial target): 26 tasks written per domain (2 spares), 24 used, so 48 in total. Spares go through the same steps.
AI answers
- Format. Code tasks: the function. Data tasks: the pandas code and the final number. Both shown in the same plain template.
- Blind to model. Checkers are never told which model produced an answer. Identifying text is removed.
- Answers per task:
- always 1 correct answer from a model, and 1 planted-error version: a correct model answer, chosen at random, with one error I inserted, so the style still matches the model;
- if the task has at least one natural error (a wrong answer a model produced on its own): 1 natural error chosen at random, plus 1 more correct answer, from the model that made the natural error if it has a correct one, otherwise from another model.
- So a task has 2 answers, or 4 if it has a natural error.
- Tasks no model solves are dropped, and the number dropped is reported. If more than 2 are dropped in a domain, the study runs with fewer tasks and reports it.
- Mix. Half of the answers are wrong. About 57% of checks are of wrong answers, because the 6-verdict subset (see Participants) is wrong answers only, as a Verification-Cost Error must be a wrong answer. Every session has at least 2 correct and at least 2 wrong answers; the allocation reports the actual distribution. Checkers are told answers "may or may not be correct", but not the proportion.
- Why balanced. Wen et al. (2024, Appendix C) also balanced their human evaluation: "We explicitly kept the balance of correct/incorrect outputs, yielding 200 examples." Limitation: real-world error rates are lower than this mix.
Planted errors
- Three types per domain:
| Code | Data analysis |
|---|---|
| Boundary or off-by-one | Wrong filter or dropped rows |
| Wrong condition | Wrong aggregation or grouping |
| Missed edge case (empty input, duplicates) | Rounding or unit |
- Two levels, checked automatically:
- Code, obvious: fails one of the worked examples in the specification, or a trivial case.
- Code, subtle: passes the worked examples and typical cases, and fails at least one hidden test.
- Data, obvious: the number is implausible on inspection (wrong order of magnitude, or an impossible value).
- Data, subtle: within 20% of the correct value, and the code runs without warnings.
- Only errors that clearly meet one of these definitions are planted.
- Realism check, pilot 2 (dress rehearsal) only. After every check, the checker answers "How likely is it that an AI wrote this exact answer?" (1–5). The aim is about 15 ratings per kind of answer (correct, planted, natural); not every answer will be rated.
- Realism threshold, decided before pilot 2 (dress rehearsal): if planted errors are rated at least 1 point lower on average than natural errors, the planted errors are rebuilt from real model errors. The question is not asked in the main study, because it could hint that some answers were altered.
Natural errors and models (RQ3)
- Models: 2–3 models of clearly different capability, with versions, temperature and prompts fixed and recorded.
- Sampling: several answers per task per model. Hidden tests or the answer key label each one right or wrong. Wrong ones are classified afterwards into the three error types for their domain, or "other", and by level: obvious, subtle, or "intermediate" if they meet neither definition. Intermediate errors are reported descriptively. When a task has several natural errors, one is chosen at random.
- Expected counts. If half the tasks have a natural error, that is 12 natural errors per domain, 24 in total, each with 3 verdicts (6 if in the subset). Spread evenly, that is about 4 per model per domain with 3 models (8 overall), or 6 per model per domain with 2 models (12 overall). The weaker model will likely supply more, and the strongest may supply very few.
- If a model makes too few errors to study, that is reported as a finding and a limitation. Planted errors are not substituted.
Conditions
- Within-subject. Each participant does a do block and a check block in one session, on different tasks.
- Six pairs per session. For each pair, the participant does one task and checks the AI's answer to the other; which is which is randomised.
- Counterbalancing. Half do first, half check first.
- No repeats. Nobody sees the same task twice, in either block.
- One domain per session. A participant may come back for the other domain in a separate session.
- One task at a time. No parallel work.
- Allowed while checking: reading, running the code, writing their own tests, a calculator. Not allowed: AI assistants or web search.
Participants
- Who: third- and fourth-year CS students and developers with up to five years' experience, with working Python and basic pandas.
- Screening: a 10-minute test (one small function, one pandas question), done online beforehand. Below a set score, not included.
- Recruitment: PES University, the Lossfunk community, and developer networks. Paid.
- Payment: a fixed fee plus a bonus for each correct answer and each correct verdict. No bonus for speed, so nobody is rewarded for rushing or for rejecting everything. The bonus may encourage guessing (the confidence ratings will show this) and using the full budget (which affects doing and checking alike).
- Session: about 90 minutes, with a break between blocks. Per session: 6 pairs, so 6 tasks done and 6 AI answers checked. Hard cap of 2 hours: pairs not finished by then are reassigned whole to another session. A half already completed is kept as extra data and shown in the allocation report. The participant is paid in full, and the overrun is recorded.
- Verdicts: 3 per answer for the primary analysis. A subset of 12 wrong answers per domain gets 6 verdicts, for an exploratory Verification-Cost Error classification. At 6 verdicts an error is confirmed only if all 6 checkers miss it, so most results will be inconclusive.
- Sample size (initial target, revised after the pilots), per domain, assuming 24 tasks (26 written, 2 spare) and half of them (12) with a natural error. If more than 2 tasks are dropped, the counts below fall and the shortfall is reported:
- Answers: 12 × 2 + 12 × 4 = 72. Wrong: 24 planted + 12 natural = 36.
- Checks: 72 × 3 = 216, plus 12 × 3 extra for the subset = 252.
- Sessions: 252 ÷ 6 = 42 per domain, 84 in total.
- Do attempts: 42 × 6 = 252, about 10 per task on average.
- Wrong checks: 36 × 3 + 12 × 3 = 144 of 252, about 57%.
- In general, with n tasks that have a natural error: 48 + 2n answers, 180 + 6n checks, 30 + n sessions.
- Fallback: if recruitment falls short, first drop the extra RQ3 answers (the natural error and its extra correct answer) and the exploratory 6-verdict subset, then use fewer tasks per domain, and only then one domain.
Allocation
Fixed by script before the main study, separately for each domain: 1. Check slots. Each answer gets as many slots as its verdict count (3, or 6 for the subset): 252 slots. 2. Pairs. A slot means a session checks that answer and does the other task in its pair. So a task is done exactly as often as its partner's answers are checked: 6 to at most 18 do attempts per task (mean about 10), depending on how many answers the partner has and how many are in the 6-verdict subset. The mixed-effects models handle unequal counts through the task random effect; tasks with fewer attempts have noisier estimates. 3. Sessions. Slots are grouped at random into sessions of 6, one per pair, so a session uses 6 different pairs and nobody sees a task twice. Every session has at least 2 correct and at least 2 wrong answers. 4. Order. Within a session, the order of pairs is random; half of the sessions do first and half check first. 5. Report. The script reports the do count per task, the verdict count per answer, how many wrong answers each session has, and any halves kept as extra data from pairs reassigned at the session cap.
Measures
- Do block: time from start to submit; correct or not against the key; whether the time cap was hit.
- Check block: time from start to verdict; verdict ("correct" or "incorrect"); confidence (50%, 75% or 100%); an optional one-line reason or counterexample; outcome, using Crescitelli et al.'s four outcomes (correct verdict, wrong verdict, budget ran out, declared unverifiable), with wrong verdicts split into false accepts and false rejects.
- Logged by the harness: number of code runs, tests written, and gaps between activity (keystroke, mouse movement, scroll or code run).
- After the session: trust in AI (1–5) and experience.
- Pilot 2 (dress rehearsal) only: after every check, how likely it is that an AI wrote that exact answer (1–5).
- Derived: verification ratio (per person and per task), time to correct verdict, false accept and false reject rates, premature closure, non-accept rate, break-even per task, and correct answers per hour for doing vs. delegating.
Protocol settings
Following Crescitelli et al. (2026), declared before the main study: - Cost unit: human minutes. - Verifiers: screened students and early-career developers. - Check budget: each task's median pilot 1 do time. The budget is "as long as it would take you to do it yourself", so a check that uses it all saves nothing. - Do-time cap: twice the task's median pilot 1 do time. - Stopping: the checker submits a verdict, the budget runs out, or they declare the answer unverifiable. - Declared deviation: Crescitelli et al. end a check only on a verdict given with "self-reported confidence ≥ τ (pre-registered, typically 0.9)" (Section 6, Step 3). Here any verdict ends the check, because forcing people to continue until they are 90% sure doesn't reflect real checking. Confidence (50%, 75% or 100%) is recorded, and premature closure is a wrong verdict given at 100% confidence. - Threshold: an error counts as a Verification-Cost Error if at least half of checkers miss it within budget. - Miss: any check that doesn't reach a correct verdict within budget: a false accept, a timeout, or "unverifiable". This follows Crescitelli et al.'s decomposition of the failure probability into "an incorrect verdict returned within budget, exhaustion of the budget without a verdict, and a declaration that the output is unverifiable" (Section 5). The breakdown into the three is reported. - Verdicts per answer: 3; 6 for a subset of 12 wrong answers per domain. - Uncertainty: 95% Clopper–Pearson intervals, with the three-state rule (confirmed, not, inconclusive). At 6 verdicts, an error is confirmed only if all 6 checkers miss it and ruled out only if none do; anything in between is inconclusive (Crescitelli et al. 2026, Section 5). - Inactivity: no automatic exclusion. Activity is a keystroke, mouse movement, scroll or code run. Gaps over 10 minutes are flagged, and results are reported with and without the flagged attempts. - Pre-registration: the full plan is registered on OSF before the main study starts.
Analysis
- Time models. Primary: a mixed-effects survival-style model of time, treating budget and cap timeouts as censored (the true time is at least the limit). Report the timeout share. Robustness check: a plain mixed-effects model of log time with the same terms.
- RQ1: terms for condition (do or check), domain, difficulty, condition × difficulty, block order and position in block, with participant and task as random effects. Report ratios with intervals, per task and per person, and the break-even per task.
- Ratios. The primary estimate of the verification ratio comes from the survival model, so timed-out attempts count as "at least this long". Raw medians of finished attempts are secondary. They leave out the slowest attempts on both sides, so they understate do and check times; which way the ratio is biased depends on whether doing or checking times out more often.
- Which attempts count. Ratios and break-even use all attempts, for both doing and checking. Correct-only versions are secondary. Time to correct verdict is reported as conditional on success.
- RQ2: mixed-effects logistic model of accepting a wrong answer, with subtlety, domain and experience as main terms. Report false accept and false reject rates with intervals. Error type is descriptive only: about 4 planted answers per type and level in each domain (24 ÷ 6).
- Reporting error rates. (a) Raw rates at the study's 50% mix. (b) The false-accept rate per wrong answer. (c) An illustration of expected outcomes per 100 delegated tasks at assumed real-world error rates of 5%, 10% and 20%, using the measured rates for wrong and correct answers. Limitation: people may check differently when errors are rarer than in the study.
- H3: compare time for false accepts with time for correct rejects, both raw and relative to each person's own do time.
- RQ3 (exploratory): per model, the natural error rate, false accept rate on its errors, and time to correct verdict, with intervals. No claims beyond description.
- Crescitelli classification (exploratory): on the 6-verdict subset, confirmed, not and inconclusive Verification-Cost Error counts, with misses broken down into false accepts, timeouts and "unverifiable". Most will be inconclusive.
- Also reported: correct answers per hour, doing vs. delegating (Ipeirotis & Zheng 2025).
Ethics and consent
- Consent: informed consent before the session; participants can stop at any time and still receive the fixed fee.
- Honesty: participants are told AI answers may be wrong. They are not told the proportion, or that some errors were inserted by hand. Afterwards they receive a debrief explaining both.
- Data: anonymous participant IDs; only data recorded inside the harness; no screen or camera recording. Released data is anonymised.
- Approval: I will seek approval from an appropriate ethics committee (PES University, or the one Lossfunk uses) before pilot 1 (timing).
- Informal walkthroughs to test the harness are labelled as such, and their data isn't used.
Hard questions
The question itself
What exactly is "verification cost" in this study?
The human minutes it takes to reach a verdict on an AI's answer within a time budget, together with whether that verdict is right. This follows Crescitelli et al. (2026). It is about human time, not machine cost.
Why compare checking with doing, instead of just measuring checking?
Checking time alone doesn't tell you whether delegating helped. Five minutes to check is good if doing the task takes thirty, and bad if it takes four. Comparing with each person's own doing time turns it into a decision.
Isn't this already known from translation research?
Partly. Translation and radiology studies have timed the same people doing a task versus correcting machine output. But there people rewrite the output, and quality is judged by humans. Here, people give a verdict only, answers are checked automatically by hidden tests, and I measure wrong answers accepted and correct ones rejected. In code, the studies I found measure checking alone (Kaufman et al. 2026) or compare different groups (Ipeirotis & Zheng 2025).
Why does this matter if models keep getting better?
Fewer errors doesn't automatically mean cheaper checking. If errors become rarer but harder to spot, checking can cost as much while more errors slip through. H4 predicts this, but it's exploratory and uncertain, and Kirchner et al. (2024) show that how checkable a model is depends on how it was trained.
Why does the break-even count timeouts and "unverifiable" as well as rejects?
Each of them leaves you without a trusted answer, so you'd redo the task. Counting only rejects would make delegation look cheaper than it is.
What result would make this project not worth continuing?
If checking is far quicker than doing and almost no wrong answers get through, on every task, then these tasks don't stress checking. I'd move to harder or longer tasks.
Tasks
Where do the tasks come from?
I write them: short Python functions, 15–40 lines, each with a written spec, two worked examples, and a hidden test suite of at least 20 tests. I draft every task myself. AI may only review specs and suggest edge-case tests. A second person solves each to confirm the spec is clear and the answer key is right.
Why not use existing benchmarks like HumanEval?
Models have probably seen them in training, so their errors wouldn't be typical, and checkers might recognise the problems.
How do you know the answer key is right?
The hidden tests are written against the spec, and a second person solves each task independently. Any disagreement means fixing the spec or the tests before the study. If it can't be resolved, the task is dropped.
Why only code, not data analysis?
With 3–4 people, two domains would halve the data for each. Code has the cleanest ground truth (hidden tests) and the clearest difference between checking and doing. Data analysis is the next phase.
Why 5–10 minute tasks?
Short enough to fit three pairs into a session, long enough for checking and doing to differ. Long, multi-step agent tasks are out of scope: METR found timing unreliable there.
How do you know a task is as hard as you labelled it?
I don't, in advance. The easy, medium and hard labels are my estimate. I check them against the observed doing times and report any mismatch.
Planted errors
Why plant errors if models make their own?
Models, especially strong ones, may not make enough errors on these tasks. Planting guarantees enough wrong answers and lets me control how subtle they are. Each task's wrong answer is a natural error where one exists, otherwise a planted one, aiming for about half of each.
Aren't planted errors unrealistic?
They might be, so I reduce and check it. Each is made by editing a correct model answer, so the style matches the model. If one targets an edge case, the spec states that case. Before the study, the walkthrough person rates a sample blind; after the study, each panel member rates answers to unseen spare tasks. If planted errors are rated at least 1 point less realistic, I report them separately and rely mainly on natural errors.
What makes an error "obvious" or "subtle"?
Obvious: it fails one of the spec's worked examples or a trivial case. Subtle: it passes the worked examples and typical cases but fails at least one hidden test. This is checked automatically. Natural errors that fit neither are labelled "intermediate".
Why only edge cases the spec already states?
Otherwise a "missed edge case" becomes an argument about what the spec meant, rather than a failure to catch an error.
Doesn't 50% wrong answers make checkers suspicious?
It might. A balanced mix gives the most information about both kinds of mistake, and Wen et al. (2024) used one too. Checkers aren't told the rate. I report the false-accept rate per wrong answer, which depends less on the mix, and note as a limitation that people may check differently when errors are rarer.
Conditions
Is it fair to compare doing one task with checking a different one?
A person never does and checks the same task, because they'd already know the answer. Each pair is matched on difficulty label and estimated doing time, and which task is done and which is checked is balanced across people with a Latin-square scheme.
Why use the same person for doing and checking?
A fast person is fast at both. Comparing within one person removes those personal differences, so each person is their own baseline.
Doesn't doing tasks teach people what errors to look for?
It could. Do-first and check-first alternate across sessions, task order is randomised, and session number is tracked in the analysis.
Won't people get better at checking over the sessions?
Probably. I track performance session by session and report it. It's a finding in itself: does checking AI get easier with practice?
What can checkers use?
They can read the code, run it, write their own tests and use a calculator. No AI assistants and no web search.
People
Why only 3–4 people?
That's what I can realistically count on from the Lossfunk community. So it's a small-N design: each person does many tasks and counts as a replication. If more people become available, for example from residency cohorts, the same design scales.
Can you say anything about programmers in general?
No. Results describe this panel. Crescitelli et al. say the same: when checkers aren't sampled to represent a population, estimates are statements about the observed panel.
What if someone drops out?
Losing one person loses a quarter to a third of the data. Sessions are short and flexible, and I screen about 5 people to get 3–4. If fewer than 3 people complete all their sessions, the human part becomes a pilot, and the instrument, task set and AI-checker comparison become the main results.
Why aren't you a participant?
I wrote the tasks and planted the errors, so I'd know the answers.
Are participants paid?
No. Panel members volunteer, with no fee, and can stop at any time. With their agreement, they're acknowledged in the write-up. If Lossfunk offers a small budget, it goes to vouchers to widen the panel.
What about differences in experience?
With 3–4 people I can't test whether experience matters. I record it and describe it. Kaufman et al. found experience didn't predict how well programmers judged LLM-written assertions.
Measurement
When does the checking clock start and stop?
It starts when the AI's answer appears and stops when the person gives a verdict, when the 10-minute budget runs out, or when they declare the answer unverifiable.
What if someone says "looks fine" in 30 seconds and is wrong?
That's a false accept, not a fast check. Time to a correct verdict only counts right verdicts, and a wrong verdict given at 100% confidence is recorded as a premature closure.
Why not make people continue until they're 90% sure, as Crescitelli suggests?
Real checking doesn't work that way. Any verdict ends the check, and confidence (50%, 75% or 100%) is recorded. This is a declared deviation from the protocol.
Why a 10-minute check budget?
It's the top of the target doing time, so a check that uses all of it saves nothing on a typical task. Doing has a 20-minute cap.
What if someone gets interrupted or stops working?
Activity means a keystroke, mouse movement, scroll or code run. Gaps over 10 minutes are flagged, and results are reported with and without them.
What if someone just guesses?
It shows up as confidence stuck at 50%, very fast verdicts, and accuracy near chance. Nobody is excluded: verdicts given in under 30 seconds at 50% confidence are flagged as possible guesses, and results are reported with and without them.
Statistics
With 3–4 people, is any of this meaningful?
For this panel, yes. Each person does 24 tasks, so each person's ratio of checking to doing time can be estimated, and each person is a replication: if all of them show the same pattern, that's evidence. For error rates it's pilot-sized, and I say so.
How do you handle checks that hit the time limit?
A survival-style model treats them as "took at least this long", rather than pretending they finished at the limit. Simple medians are reported as a secondary check, with their bias stated.
How is the ratio of checking to doing time calculated?
The main estimate comes from the survival model. Per person, intervals come from resampling their tasks. Pooled across people, person is a fixed effect and task a random effect.
Why not classify Verification-Cost Errors, as Crescitelli proposes?
That needs at least 6 checkers per answer, and even then an error is only confirmed if all 6 miss it (Crescitelli et al., Section 5). Here each answer gets at most 1 human verdict, so I report per answer whether its checker caught it, and error rates are pooled across answers.
How many wrong answers will actually be checked?
About 18 with 3 people, about 24 with 4. That gives a false-accept rate with a wide interval, which I state plainly.
AI models
Which models?
One strong closed model and one or two open-weights models of different sizes, with exact versions and settings recorded.
Could the models cheat or break out of a sandbox?
They never take part in the live study. Each model answers each task 5 times, offline, and the answers are saved as text. The saved code is run against the hidden tests on my machine, in a container with no internet and a time limit. The hidden tests are never seen by models or sent to participants' browsers. The real risk is that they saw similar tasks in training, which is why the tasks are new.
What if a model gets every task right?
Then it makes no natural errors here, and I report that as a finding. Planted errors still provide the wrong answers for the main questions.
Why use AI checkers if the project is about human cost?
They're a supporting part. They ask whether an AI checking first could reduce human checking cost. They're cheap, they scale where a 3–4 person panel can't, and they keep the project useful if fewer people are available. They don't measure verification cost themselves.
Why can't the AI checkers run code, when humans can?
To keep the first version simple: text-only judgement. That makes it an unequal comparison, and I say so. A version where AI checkers can run code is a natural extension.
Will a model ever judge its own answers?
No.
Ethics
Do you need ethics approval?
Yes, before any data is collected from people.
Which ethics review will you use?
Don't know yet — how I'd find out: ask Lossfunk what review process they use, and otherwise apply to PES University's ethics committee.
Isn't hiding planted errors deception?
Participants are told the AI's answers may be wrong, which is true. They aren't told the proportion, or that some errors were inserted by hand. A debrief afterwards explains both. The ethics review decides whether that is acceptable.
What data do you keep?
Anonymous IDs. What the harness records, plus screening results and realism ratings, stored under the same IDs. No screen or camera recording. Any released data is anonymised.
Feasibility
Can you do this in six months?
Months 1–2: tasks, harness and ethics. Months 3–4: panel sessions, about 6–8 weeks. Months 5–6: one extension, chosen at the end of April, from: a data-analysis domain pilot; AI checkers that can run code; re-running the newest models; or more tasks with the same panel (16 new tasks), as a separately pre-registered follow-up. Then releasing the tool and data, and writing up. The core fits comfortably; the extension is planned, not improvised.
How much will this cost Lossfunk?
Very little. Hosting is free. Participants' own code runs in their browser, and grading happens offline on my machine, so there's no server running anyone's code. Model answers and AI checkers are expected to cost a few dollars, capped at US$50. The real cost is the panel members' time: about 3 hours each.
Is writing 32 tasks realistic?
About 2–3 weeks, including hidden tests and second-person checks, in month 1.
What's the biggest risk?
Not having enough people. I state the assumption openly and raise it early. The project produces the instrument, the task set and the AI-checker comparison regardless.
What will exist at the end?
An open harness, the task set with planted and natural errors and hidden tests (except the held-out set: the spares used for the realism check, whose task text stays private), anonymised raw data, the realism ratings with task IDs and error types, the AI-checker comparison, and a write-up.
Interpretation
What result would surprise you?
Checking taking as long as doing even on easy tasks; or subtle errors almost never getting through; or humans and AI checkers missing completely different errors.
If checking turns out cheaper than doing, is that the end of it?
No. Wrong answers that get through are a quality cost. Faster isn't a win if more wrong answers get through: Ipeirotis & Zheng found an AI tool made people faster without more correct answers per hour.
Can you say stronger models are harder to check?
Not as a causal claim. With a few errors per model, I can only describe what happened: on these tasks, a given model's wrong answers were caught more or less often.
How could someone else build on this?
Use the open harness with more people, other kinds of tasks or newer models. Because the protocol is declared in advance, their results can be compared with these.
Six months
Overview
The fellowship runs January to June 2027, my final semester, which is a full-time internship semester. I'd be on-site full-time for all six months.
The plan has four phases:
| Months | Phase | What it produces |
|---|---|---|
| January | Build the tasks | 32 checked tasks, model answers, graded |
| February | Build the instrument | Working harness, planted errors, allocation, AI-checker results, pre-registration |
| March–April | Run the panel | Human data from 3–4 people |
| May–June | Analyse, extend, release | Results, one extension, open release, write-up |
The core study fits comfortably in four months. The last two go into analysis, one planned extension, and releasing everything openly.
Assumptions
- 3–4 people from the Lossfunk community can each give about 3 hours, in March and April. This is the first thing I'd confirm.
- Ethics approval can be obtained by the end of February. It is needed before any data is collected from people.
- The models I use stay available long enough to generate and save all their answers in January.
What success looks like
| Level | What exists at the end |
|---|---|
| Minimum | The open harness, the checked task set and the AI-checker comparison, released, none of which depend on the panel; plus a pilot-sized human dataset if at least some people take part. |
| Target | The full core study with 3–4 people (ideally 4): verification ratios per person, false-accept rates, H1–H3 tested descriptively, RQ3 and RQ4 described, released with a write-up. |
| Stretch | The target, plus one of the May extensions, plus a workshop paper or arXiv preprint. |
The minimum is reachable even if recruitment fails, because the instrument, the tasks and the AI-checker comparison don't depend on the panel.
Milestones
| Month | Goal | Deliverable |
|---|---|---|
| January | Tasks written, checked and answered by models | 32 tasks (24 used, 6 spare, 2 practice) with specs, worked examples and hidden tests; model answers graded |
| February | Instrument ready | Harness, planted errors, pairs and allocation, ethics approval, OSF pre-registration, then AI-checker results v1; panel provisionally confirmed |
| March | Panel, first half | Screening and the walkthrough person's realism ratings (week 9); practice session and sessions 1–2 for every person; weekly data checks |
| April | Panel, second half | Sessions 3–4 (and a 5th if needed), realism ratings, debriefs; complete dataset |
| May | Analysis and one extension | Full analysis as pre-registered; human vs. AI-checker comparison; extension started |
| June | Release and write-up | Open release of harness, all 24 used tasks, anonymised data and analysis code (the held-out set, the spares used for the realism check, keeps its task text private; its realism ratings are released with task IDs and error types); write-up; talk at Lossfunk |
Month by month
January: build the tasks
- Week 1
- Confirm with Lossfunk who can join the panel, and when (checkpoint 0).
- Find out which ethics review to use, and submit the application.
- Choose the models (one strong closed model, one or two open-weights models of different sizes) and record exact versions and settings.
- Set up the task format and the edge-case checklist (empty input, duplicates, ties, negative values, boundary values).
- Weeks 2–3
- Write the 32 tasks, about 11 a week. Each has a written spec, two worked examples, and at least 20 hidden tests.
- Label each task easy, medium or hard (8 of each among the 24 used).
- AI may only review specs and suggest edge-case tests; I draft every task myself.
- Week 4
- A second person solves every task to check the spec and the answer key. Fix or drop any task that fails; drops come out of the spares.
- Each model answers each task 5 times, offline. Run every answer against the hidden tests in a sandbox (a container with no internet and a time limit).
- End of January: a checked task set, and graded model answers showing which tasks have natural errors.
February: build the instrument
- Week 5
- Choose each task's wrong answer: a natural error where one exists, otherwise a planted one. Aim for about half of each.
- Plant errors only on edge cases the spec states, at two levels (obvious and subtle), and check each level automatically.
- Pair tasks within each difficulty label, and generate the balanced Latin-square allocation.
- Week 6
- Build the harness: one task at a time; participants run their own code in the browser; timing, verdict and confidence; activity logging.
- Build the offline grading pipeline. Hidden tests never go to the browser.
- Week 7
- Informal walkthrough with one person, to test the software only. No timing or verdict data is kept.
- Fix everything the walkthrough finds.
- Week 8
- Pre-register the core study on OSF, including the 2-pair contingency.
- Then run the AI checkers: 3–5 models judge every answer 5 times, text-only, never their own answers. This gives results before any human session.
- End of February (checkpoint 1): ethics approved, pre-registration done, harness working, at least 3 people provisionally confirmed. Screening follows in week 9.
March: panel, first half
- Week 9 (after ethics approval)
- Run the screening test with about 5 people, and schedule sessions.
- In a separate, consented step, the walkthrough person rates a sample of answers for realism.
- From week 10: each person does a practice session, then sessions 1 and 2. Sessions are about 45 minutes, with a hard stop at 60.
- Reserve rule: if someone drops out before their first session, a reserve from screening takes their place in the allocation. After their first session, no replacements.
- Weekly checks on data quality only (logging, grading, timing gaps), not on results.
- Mid-March (checkpoint 2): do real sessions fit in 60 minutes? If they often run over, use the pre-registered 2-pair contingency.
April: panel, second half
- Sessions 3 and 4, and a 5th if anything was left over.
- After each person's final session: realism ratings on spare tasks, then a debrief explaining the planted errors.
- Grade all submitted answers offline.
- End of April (checkpoint 3): how many people completed all their sessions?
May: analyse and extend
- Run the pre-registered analysis:
- each person's verification ratio, and by difficulty (H1);
- false-accept rates by subtlety (H2);
- wrong accepts versus correct rejects (H3);
- each model's natural errors (RQ3);
- human versus AI-checker misses (RQ4);
- session trends, with and without flagged guesses.
- Start one extension, chosen at checkpoint 3:
- a data-analysis domain pilot (the design already exists);
- AI checkers that can run code;
- re-running the newest models, to add a second point in time;
- more tasks with the same panel (16 new tasks), as a separately pre-registered follow-up.
- If the final results show the tasks were too easy, the extension uses harder tasks.
June: release and write up
- Release the harness, all 24 used tasks with the data, the anonymised data, and the analysis code. The held-out set is the spares used for the realism check, never tasks used in the study. Their task text stays private, so future models can't simply learn it; their realism ratings are released with task IDs and error types.
- Write up the results honestly at their real size.
- Give a talk at Lossfunk.
- Write a short guide so anyone can run the study with more people.
Change direction if…
| When | Signal | What I'd do |
|---|---|---|
| January, week 1 | Lossfunk can't confirm any panel members | Treat the instrument, the tasks and the AI-checker comparison as the main results; look for volunteers from residency cohorts |
| End of January | Tasks are taking much longer than planned | Write 24 + 4 spares instead of 6; never cut the second-person check. With only 4 spares, any drop may push the realism check onto tasks people did |
| End of February | Ethics approval not granted yet | Keep building, and run the AI checkers. If it isn't granted by mid-March, the human part shrinks to a pilot |
| February, walkthrough | Tasks take far longer or shorter than 5–10 minutes | Rewrite them using the spares before any panel session |
| February, before the panel | Walkthrough timings or model error rates suggest tasks are too easy | Rewrite affected tasks using the spares before any panel session |
| Mid-March | Sessions often run past 60 minutes | Move to 2 pairs per session and add sessions. This is a contingency declared in the pre-registration; the pairs and allocation don't change |
| End of April | Fewer than 3 people completed all their sessions | The human part becomes a pilot; the instrument, the tasks and the AI-checker comparison become the main results |
| January–April | Strong models make almost no natural errors | Use planted errors for the main questions, and report RQ3 as a limitation. Never relabel planted errors as a model comparison |
| April | Planted errors rated at least 1 point less realistic | Report them separately and base conclusions mainly on natural errors |
Risks
| Risk | Likelihood | What I'd do |
|---|---|---|
| Not enough people for the panel | High | Confirm in week 1; screen 5 to get 3–4; the minimum success level doesn't depend on the panel |
| Someone drops out partway | Medium | Short, flexible sessions; a reserve from screening only before their first session; report what was completed |
| Ethics approval is slow | Medium | Apply in week 1; build everything else meanwhile |
| Writing 32 good tasks takes longer | Medium | About 11 a week; reduce spares if needed, never the checks |
| Harness bugs or timing errors | Medium | Walkthrough first; keep raw event logs; weekly data checks |
| Hidden tests leak | Low | Never sent to the browser; grading only offline |
| Models make too few natural errors | Medium | Planted errors cover the main questions; RQ3 becomes descriptive or a limitation |
| Planted errors look fake | Medium | Made from real model answers; realism ratings; report separately if needed |
| People improve with practice | High | Expected, tracked session by session, and reported as a finding |
| A model is retired or changes | Low | Exact versions recorded; all answers saved in January |
| Results are small and uncertain | High | Stated plainly as pilot-sized; the instrument is the lasting output |
| Released tasks end up in future training data | Medium | Keep the held-out set private: the spares used for the realism check |
What I need from Lossfunk
- People: 3–4 community members, about 3 hours each, in March and April (about 5 hours each if possible).
- Ethics: advice on which review process to use.
- A small budget for API calls: a few dollars expected, capped at US$50, for model answers and AI checkers.
- A mentor: a short check-in every week or two.
- A quiet room for sessions.
- Optional: a small voucher budget, or access to residency cohorts, to widen the panel.
Budget
| Item | Cost |
|---|---|
| Participants | Volunteers, no fee |
| Hosting the harness | Free tier |
| Running participants' code | Free: it runs in their own browser |
| Model answers and AI checkers | A few dollars expected, capped at US$50 |
| Grading | Free: offline on my machine |
How I'll work
- Weekly: a plan on Monday and a short written update on Friday, kept in the repo.
- Every week or two: a check-in with my mentor.
- Everything versioned: tasks, code, allocation and analysis in git, so any result can be traced to the exact version that produced it.
- No peeking: during the panel, I check data quality only, not results, so the pre-registered analysis stays honest.
Before January (only if accepted)
I'll be finishing my seventh semester, so this is light and optional: - read more deeply on the key papers; - draft a few practice tasks; - find out which ethics review applies, so I can apply in week 1; - find out what PES needs for the internship to count (offer letter, any reports).
Pilot
Status
No pilot has been run yet. The design requires ethics approval before any data is collected from people, and the panel comes from the Lossfunk community.
Planned
A software-only walkthrough of the harness (February, week 7), then screening and realism ratings after ethics approval (March, week 9). See PLAN.md.
Already tested
The spot-the-bug puzzle on this page, whose bug was verified with tests.
Working with AI
Caught by me
- I challenged the gap claim. The AI's draft said nobody had timed people doing versus checking AI work. I asked for a proper literature review before accepting it. It found translation and radiology studies had done this, so I narrowed the claim to code with exact answers. The first draft had claimed nobody had measured this at all; METR, Wen et al. and Tufano et al. had measured parts of it.
- I rejected a design I couldn't run. The design at that point needed 84 sessions, far more people than I could realistically recruit. I could count on 3–4 people, so the study was redesigned as a small-N panel, with the full design kept as the "if resourced" version.
Caught by checking the AI's work
- A design review found three flaws in the draft PROBLEM.md. Showing only a number for data-analysis answers would make checking the same as redoing; the break-even ignored checks that time out; H3 overclaimed a replication.
- A cited paper contradicted my own hypothesis. Kaufman et al. found experience did not predict checking accuracy, which H2 assumed. I made it an open question.
- A statistics claim was wrong. The draft said six checkers per answer were the minimum for confident classification. Under the protocol's own rule, six only confirms an error when all six miss it.
- A citation didn't hold up. A balanced mix of right and wrong answers was credited to Kirchner et al.; verification couldn't confirm it for human checkers, so I cited Wen et al., who state it explicitly.
- The design would have leaked the answers. Grading in the browser would have exposed the hidden tests to anyone using developer tools. Grading now happens offline.
- Some summaries could mislead. Checking each number against the papers found that METR's 9% was time spent reviewing and cleaning up AI output, and that Tufano et al.'s "no time saved" hid that manual reviews were, if anything, faster. Both rows were reworded.