Measuring the Human Cost of Verifying AI Output

In one line

When you could do a task yourself, does checking an AI's attempt cost less time than doing it? And how often does a wrong answer get through?

Links

Status

Proposal for the Lossfunk Fellowship, October 2026. Not yet run: the study needs ethics approval and a panel first.

Why I can run this

The problem

In one sentence

When a person could do a task themselves, does checking an AI's attempt cost less of their time than doing it, and how often does a wrong AI answer get through the check?

The problem

Delegating work to AI assumes that checking its output is cheaper than doing the work. If checking takes about as long as doing, or wrong answers slip through, delegation saves little or nothing.

Other fields have measured this. Machine-translation research has timed the same translators translating from scratch versus correcting machine output (Plitt & Masselot 2010; Green et al. 2013; Sarti et al. 2022), and a 2024 radiology pilot did the same with deliberately planted AI errors (Acosta et al. 2024). But in those studies people rewrite the output, and quality is judged by humans, not against an exact answer.

In code and data analysis, the studies I found measure only part of it: - Checking alone. Programmers judging LLM-written assertions accepted about half of the wrong ones, and decided on wrong accepts faster than on correct rejects (Kaufman et al. 2026). - Doing vs. AI-assisted work, in different groups. SQL users with an AI tool finished about 31% faster than users writing by hand, but produced roughly the same number of correct answers per hour, and were somewhat less accurate (46% vs. 64% of queries correct, not statistically significant) (Ipeirotis & Zheng 2025). - With vs. without AI, end to end. Developers were 19% slower with AI while believing they were faster (METR 2025); AI code reviews saved no time (Tufano et al. 2025).

Crescitelli et al. (2026) proposed a protocol for measuring verification cost against a time budget, but ran no experiment.

The gap: in this search I did not find a study in code or data analysis that times the same people solving tasks and judging AI attempts at matched tasks, with automatically checkable answers, and reports check time, do time, and the rates of wrong answers accepted and correct answers rejected.

Key definitions

Scope and setting

Research questions

  1. Time. On matched tasks, how does check time compare with do time, and how does the verification ratio change with difficulty? Where does delegation fail to break even?
  2. Errors that get through. How often are wrong AI answers accepted within budget, and how does that depend on subtlety? In the core study this is descriptive: acceptance by subtlety and by error type, with counts.
  3. Across models (exploratory). For errors the models make on their own, are a stronger model's wrong answers caught less often, or more slowly, than a weaker model's?
  4. AI checkers (supporting). Do AI checkers miss the same errors that human checkers miss?

Hypotheses

What would change my mind

Out of scope

References

What's known

Papers

Doing vs. correcting AI output, in other fields

The same people do a task from scratch and correct machine output, both timed.

Paper What they did What they left open
Plitt & Masselot 2010 (PBML 93) PDF 12 professional translators from Autodesk's localisation vendors, timed translating from scratch vs. post-editing machine translation. Post-editing saved about 43% of translation time. People rewrite rather than judge. Quality rated by humans, not an exact key. No measure of errors let through.
Green, Heer & Manning 2013 (CHI, DOI) 48 professional translators, 3 languages. The same people translated and post-edited, on different sentences. Post-editing was faster and produced better quality. Same as above.
Sarti et al. 2022 (DivEMT, EMNLP) ACL Anthology 18 professional translators, 6 languages, keystroke and timing logs. Speed-ups from post-editing ranged from about 10% to about double, depending on language and translation system. Same as above. Shows the benefit depends heavily on output quality.
Acosta et al. 2024 (arXiv 2412.12042) 3 radiology readers (one radiologist, two residents) wrote reports from a template or edited AI drafts, with errors planted in half the drafts. Median reporting time fell from 573 to 435 seconds. Tiny sample. Expert-judged, not an exact key. Editing, not a verdict.

Checking AI output only

People judge AI output, but never do the task themselves.

Paper What they did What they left open
Kaufman et al. 2026 (arXiv 2607.08885) 86 Python programmers judged LLM-written assertions. They accepted about half of the wrong ones, and decided on wrong accepts faster than on correct rejects. No do-it-yourself comparison. Judged small assertions, not full solutions.
Wen et al. 2024 (arXiv 2409.12822) Time-limited checkers (3–10 min) judged answers from one model before and after RLHF. Accepting wrong answers rose by 24.1 points on reading questions and 18.3 on code. No do-it-yourself comparison. One model, before vs. after training.
Kirchner et al. 2024 (arXiv 2407.13692) Checkers had 45 seconds per maths solution. Training for correctness only made solutions slower and harder to check; training for checkability reversed that. No do-it-yourself comparison. Easy school maths, very short budget.
Mozannar et al. 2024 (CHI, arXiv 2210.14306) 21 programmers using Copilot labelled their own screen recordings. Verifying suggestions took 22.4% of their time. Autocomplete only. No record of wrong code accepted.

With vs. without AI, end to end

Total time is compared, but checking isn't separated out.

Paper What they did What they left open
Becker et al. 2025 (METR, arXiv 2507.09089) 16 experienced developers, 246 real tasks randomly allowed or disallowed AI. They were 19% slower with AI while believing they were faster. About 9% of AI time went to reviewing and cleaning up AI output. Total time only. No separate checking step. No record of wrong output accepted.
Tufano et al. 2025 (ICSE, arXiv 2411.11401) 29 professional developers reviewed code manually, with a ChatGPT review, or with a secretly "perfect" review. Neither AI condition saved time; manual reviews were, if anything, faster, though the differences weren't significant; AI reviews made people focus only where the AI pointed. Reviewers found a median of about half the injected issues, with or without the ChatGPT review. Review comments, not verdicts on solutions. One model, small sample.
Ipeirotis & Zheng 2025 (arXiv 2511.14718, v3 2026) 20 SQL users on tasks with known answers: half wrote by hand, half used an AI tool. The AI group was about 31% faster, but produced roughly the same number of correct answers per hour, and were somewhat less accurate (46% vs. 64%, not statistically significant). Different people in each group. No deliberately wrong queries.

Protocol and theory

Paper What they did What they left open
Crescitelli et al. 2026 (arXiv 2608.08709) Defined Verification-Cost Errors (wrong answers enough checkers miss within budget) and a six-step protocol to measure them. Never run. No do-it-yourself comparison.
Vasconcelos et al. 2023 (CSCW, arXiv 2212.06823) 731 people across 5 studies. People over-rely on AI more when checking it is costly. The theory behind the question, not a do-vs-check measurement.

The gap

Other fields have timed the same people doing a task versus correcting machine output, and code studies have measured checking alone or doing versus AI-assisted work in separate groups. In this search I did not find a study in code or data analysis that times the same people solving tasks and judging AI attempts at matched tasks, with automatically checkable answers, and reports wrong answers accepted and correct answers rejected.

What I borrow

The study

Core study (what I'll run)

Overview

A panel of 3–4 people from the Lossfunk community each solve 12 code tasks themselves and judge AI answers to 12 matched tasks. AI models also judge every AI answer, as a supporting comparison.

flowchart TD
  A["Write 32 code tasks (24 used, 6 spare, 2 practice)"] --> B[Model answers: 1 correct and 1 wrong per task]
  B --> C[Pair tasks]
  C --> D[Do: solve one task of each pair]
  C --> E[Check: judge the AI answer to its partner]
  P["Panel: 3–4 people, sessions of 3 pairs (max 60 min)"] --> D
  P --> E
  D --> G[Do time and correctness]
  E --> H[Check time, verdict and confidence]
  G --> I[Per-person ratios and false accepts]
  H --> I
  B --> J["AI checkers: 3–5 models judge every answer 5 times"]
  H --> K[Errors missed by humans vs. AI checkers]
  J --> K

Domain

Panel

Tasks and time

AI answers

Where code runs

Allocation

Timing (no pilot participants available)

Realism

Analysis

AI checkers (supporting part)

Pre-registration

Ethics, payment and consent

Full study (if resourced)

The design for a larger study, kept for when more participants are available. Where it differs, the core study above takes precedence.

Overview

Participants solve some tasks themselves and judge AI answers to matched tasks. Everything runs in a web harness that shows one task at a time and records timing, code runs and verdicts.

Order of steps

  1. Ethics approval.
  2. Write 26 tasks per domain, with specs written against the edge-case checklist (see Tasks and domains); a second person checks each.
  3. Generate model answers.
  4. Drop tasks no model solves; keep 24 per domain. If 0 or 1 tasks are dropped, the 24 kept are chosen at random, with the random seed recorded (do times aren't known until pilot 1, so they can't be used).
  5. Plant errors.
  6. Pilot 1 (timing): set budgets and difficulty.
  7. Pair tasks.
  8. Pilot 2 (dress rehearsal): test harness, allocation, session length; realism ratings.
  9. Finalise allocation.
  10. Pre-register on OSF.
  11. Main study.
  12. Analysis.

Pairing comes after dropping, so a dropped task never breaks a pair.

Tasks and domains

AI answers

Planted errors

Code Data analysis
Boundary or off-by-one Wrong filter or dropped rows
Wrong condition Wrong aggregation or grouping
Missed edge case (empty input, duplicates) Rounding or unit

Natural errors and models (RQ3)

Conditions

Participants

Allocation

Fixed by script before the main study, separately for each domain: 1. Check slots. Each answer gets as many slots as its verdict count (3, or 6 for the subset): 252 slots. 2. Pairs. A slot means a session checks that answer and does the other task in its pair. So a task is done exactly as often as its partner's answers are checked: 6 to at most 18 do attempts per task (mean about 10), depending on how many answers the partner has and how many are in the 6-verdict subset. The mixed-effects models handle unequal counts through the task random effect; tasks with fewer attempts have noisier estimates. 3. Sessions. Slots are grouped at random into sessions of 6, one per pair, so a session uses 6 different pairs and nobody sees a task twice. Every session has at least 2 correct and at least 2 wrong answers. 4. Order. Within a session, the order of pairs is random; half of the sessions do first and half check first. 5. Report. The script reports the do count per task, the verdict count per answer, how many wrong answers each session has, and any halves kept as extra data from pairs reassigned at the session cap.

Measures

Protocol settings

Following Crescitelli et al. (2026), declared before the main study: - Cost unit: human minutes. - Verifiers: screened students and early-career developers. - Check budget: each task's median pilot 1 do time. The budget is "as long as it would take you to do it yourself", so a check that uses it all saves nothing. - Do-time cap: twice the task's median pilot 1 do time. - Stopping: the checker submits a verdict, the budget runs out, or they declare the answer unverifiable. - Declared deviation: Crescitelli et al. end a check only on a verdict given with "self-reported confidence ≥ τ (pre-registered, typically 0.9)" (Section 6, Step 3). Here any verdict ends the check, because forcing people to continue until they are 90% sure doesn't reflect real checking. Confidence (50%, 75% or 100%) is recorded, and premature closure is a wrong verdict given at 100% confidence. - Threshold: an error counts as a Verification-Cost Error if at least half of checkers miss it within budget. - Miss: any check that doesn't reach a correct verdict within budget: a false accept, a timeout, or "unverifiable". This follows Crescitelli et al.'s decomposition of the failure probability into "an incorrect verdict returned within budget, exhaustion of the budget without a verdict, and a declaration that the output is unverifiable" (Section 5). The breakdown into the three is reported. - Verdicts per answer: 3; 6 for a subset of 12 wrong answers per domain. - Uncertainty: 95% Clopper–Pearson intervals, with the three-state rule (confirmed, not, inconclusive). At 6 verdicts, an error is confirmed only if all 6 checkers miss it and ruled out only if none do; anything in between is inconclusive (Crescitelli et al. 2026, Section 5). - Inactivity: no automatic exclusion. Activity is a keystroke, mouse movement, scroll or code run. Gaps over 10 minutes are flagged, and results are reported with and without the flagged attempts. - Pre-registration: the full plan is registered on OSF before the main study starts.

Analysis

Ethics and consent

Hard questions

The question itself

What exactly is "verification cost" in this study?

The human minutes it takes to reach a verdict on an AI's answer within a time budget, together with whether that verdict is right. This follows Crescitelli et al. (2026). It is about human time, not machine cost.

Why compare checking with doing, instead of just measuring checking?

Checking time alone doesn't tell you whether delegating helped. Five minutes to check is good if doing the task takes thirty, and bad if it takes four. Comparing with each person's own doing time turns it into a decision.

Isn't this already known from translation research?

Partly. Translation and radiology studies have timed the same people doing a task versus correcting machine output. But there people rewrite the output, and quality is judged by humans. Here, people give a verdict only, answers are checked automatically by hidden tests, and I measure wrong answers accepted and correct ones rejected. In code, the studies I found measure checking alone (Kaufman et al. 2026) or compare different groups (Ipeirotis & Zheng 2025).

Why does this matter if models keep getting better?

Fewer errors doesn't automatically mean cheaper checking. If errors become rarer but harder to spot, checking can cost as much while more errors slip through. H4 predicts this, but it's exploratory and uncertain, and Kirchner et al. (2024) show that how checkable a model is depends on how it was trained.

Why does the break-even count timeouts and "unverifiable" as well as rejects?

Each of them leaves you without a trusted answer, so you'd redo the task. Counting only rejects would make delegation look cheaper than it is.

What result would make this project not worth continuing?

If checking is far quicker than doing and almost no wrong answers get through, on every task, then these tasks don't stress checking. I'd move to harder or longer tasks.

Tasks

Where do the tasks come from?

I write them: short Python functions, 15–40 lines, each with a written spec, two worked examples, and a hidden test suite of at least 20 tests. I draft every task myself. AI may only review specs and suggest edge-case tests. A second person solves each to confirm the spec is clear and the answer key is right.

Why not use existing benchmarks like HumanEval?

Models have probably seen them in training, so their errors wouldn't be typical, and checkers might recognise the problems.

How do you know the answer key is right?

The hidden tests are written against the spec, and a second person solves each task independently. Any disagreement means fixing the spec or the tests before the study. If it can't be resolved, the task is dropped.

Why only code, not data analysis?

With 3–4 people, two domains would halve the data for each. Code has the cleanest ground truth (hidden tests) and the clearest difference between checking and doing. Data analysis is the next phase.

Why 5–10 minute tasks?

Short enough to fit three pairs into a session, long enough for checking and doing to differ. Long, multi-step agent tasks are out of scope: METR found timing unreliable there.

How do you know a task is as hard as you labelled it?

I don't, in advance. The easy, medium and hard labels are my estimate. I check them against the observed doing times and report any mismatch.

Planted errors

Why plant errors if models make their own?

Models, especially strong ones, may not make enough errors on these tasks. Planting guarantees enough wrong answers and lets me control how subtle they are. Each task's wrong answer is a natural error where one exists, otherwise a planted one, aiming for about half of each.

Aren't planted errors unrealistic?

They might be, so I reduce and check it. Each is made by editing a correct model answer, so the style matches the model. If one targets an edge case, the spec states that case. Before the study, the walkthrough person rates a sample blind; after the study, each panel member rates answers to unseen spare tasks. If planted errors are rated at least 1 point less realistic, I report them separately and rely mainly on natural errors.

What makes an error "obvious" or "subtle"?

Obvious: it fails one of the spec's worked examples or a trivial case. Subtle: it passes the worked examples and typical cases but fails at least one hidden test. This is checked automatically. Natural errors that fit neither are labelled "intermediate".

Why only edge cases the spec already states?

Otherwise a "missed edge case" becomes an argument about what the spec meant, rather than a failure to catch an error.

Doesn't 50% wrong answers make checkers suspicious?

It might. A balanced mix gives the most information about both kinds of mistake, and Wen et al. (2024) used one too. Checkers aren't told the rate. I report the false-accept rate per wrong answer, which depends less on the mix, and note as a limitation that people may check differently when errors are rarer.

Conditions

Is it fair to compare doing one task with checking a different one?

A person never does and checks the same task, because they'd already know the answer. Each pair is matched on difficulty label and estimated doing time, and which task is done and which is checked is balanced across people with a Latin-square scheme.

Why use the same person for doing and checking?

A fast person is fast at both. Comparing within one person removes those personal differences, so each person is their own baseline.

Doesn't doing tasks teach people what errors to look for?

It could. Do-first and check-first alternate across sessions, task order is randomised, and session number is tracked in the analysis.

Won't people get better at checking over the sessions?

Probably. I track performance session by session and report it. It's a finding in itself: does checking AI get easier with practice?

What can checkers use?

They can read the code, run it, write their own tests and use a calculator. No AI assistants and no web search.

People

Why only 3–4 people?

That's what I can realistically count on from the Lossfunk community. So it's a small-N design: each person does many tasks and counts as a replication. If more people become available, for example from residency cohorts, the same design scales.

Can you say anything about programmers in general?

No. Results describe this panel. Crescitelli et al. say the same: when checkers aren't sampled to represent a population, estimates are statements about the observed panel.

What if someone drops out?

Losing one person loses a quarter to a third of the data. Sessions are short and flexible, and I screen about 5 people to get 3–4. If fewer than 3 people complete all their sessions, the human part becomes a pilot, and the instrument, task set and AI-checker comparison become the main results.

Why aren't you a participant?

I wrote the tasks and planted the errors, so I'd know the answers.

Are participants paid?

No. Panel members volunteer, with no fee, and can stop at any time. With their agreement, they're acknowledged in the write-up. If Lossfunk offers a small budget, it goes to vouchers to widen the panel.

What about differences in experience?

With 3–4 people I can't test whether experience matters. I record it and describe it. Kaufman et al. found experience didn't predict how well programmers judged LLM-written assertions.

Measurement

When does the checking clock start and stop?

It starts when the AI's answer appears and stops when the person gives a verdict, when the 10-minute budget runs out, or when they declare the answer unverifiable.

What if someone says "looks fine" in 30 seconds and is wrong?

That's a false accept, not a fast check. Time to a correct verdict only counts right verdicts, and a wrong verdict given at 100% confidence is recorded as a premature closure.

Why not make people continue until they're 90% sure, as Crescitelli suggests?

Real checking doesn't work that way. Any verdict ends the check, and confidence (50%, 75% or 100%) is recorded. This is a declared deviation from the protocol.

Why a 10-minute check budget?

It's the top of the target doing time, so a check that uses all of it saves nothing on a typical task. Doing has a 20-minute cap.

What if someone gets interrupted or stops working?

Activity means a keystroke, mouse movement, scroll or code run. Gaps over 10 minutes are flagged, and results are reported with and without them.

What if someone just guesses?

It shows up as confidence stuck at 50%, very fast verdicts, and accuracy near chance. Nobody is excluded: verdicts given in under 30 seconds at 50% confidence are flagged as possible guesses, and results are reported with and without them.

Statistics

With 3–4 people, is any of this meaningful?

For this panel, yes. Each person does 24 tasks, so each person's ratio of checking to doing time can be estimated, and each person is a replication: if all of them show the same pattern, that's evidence. For error rates it's pilot-sized, and I say so.

How do you handle checks that hit the time limit?

A survival-style model treats them as "took at least this long", rather than pretending they finished at the limit. Simple medians are reported as a secondary check, with their bias stated.

How is the ratio of checking to doing time calculated?

The main estimate comes from the survival model. Per person, intervals come from resampling their tasks. Pooled across people, person is a fixed effect and task a random effect.

Why not classify Verification-Cost Errors, as Crescitelli proposes?

That needs at least 6 checkers per answer, and even then an error is only confirmed if all 6 miss it (Crescitelli et al., Section 5). Here each answer gets at most 1 human verdict, so I report per answer whether its checker caught it, and error rates are pooled across answers.

How many wrong answers will actually be checked?

About 18 with 3 people, about 24 with 4. That gives a false-accept rate with a wide interval, which I state plainly.

AI models

Which models?

One strong closed model and one or two open-weights models of different sizes, with exact versions and settings recorded.

Could the models cheat or break out of a sandbox?

They never take part in the live study. Each model answers each task 5 times, offline, and the answers are saved as text. The saved code is run against the hidden tests on my machine, in a container with no internet and a time limit. The hidden tests are never seen by models or sent to participants' browsers. The real risk is that they saw similar tasks in training, which is why the tasks are new.

What if a model gets every task right?

Then it makes no natural errors here, and I report that as a finding. Planted errors still provide the wrong answers for the main questions.

Why use AI checkers if the project is about human cost?

They're a supporting part. They ask whether an AI checking first could reduce human checking cost. They're cheap, they scale where a 3–4 person panel can't, and they keep the project useful if fewer people are available. They don't measure verification cost themselves.

Why can't the AI checkers run code, when humans can?

To keep the first version simple: text-only judgement. That makes it an unequal comparison, and I say so. A version where AI checkers can run code is a natural extension.

Will a model ever judge its own answers?

No.

Ethics

Do you need ethics approval?

Yes, before any data is collected from people.

Which ethics review will you use?

Don't know yet — how I'd find out: ask Lossfunk what review process they use, and otherwise apply to PES University's ethics committee.

Isn't hiding planted errors deception?

Participants are told the AI's answers may be wrong, which is true. They aren't told the proportion, or that some errors were inserted by hand. A debrief afterwards explains both. The ethics review decides whether that is acceptable.

What data do you keep?

Anonymous IDs. What the harness records, plus screening results and realism ratings, stored under the same IDs. No screen or camera recording. Any released data is anonymised.

Feasibility

Can you do this in six months?

Months 1–2: tasks, harness and ethics. Months 3–4: panel sessions, about 6–8 weeks. Months 5–6: one extension, chosen at the end of April, from: a data-analysis domain pilot; AI checkers that can run code; re-running the newest models; or more tasks with the same panel (16 new tasks), as a separately pre-registered follow-up. Then releasing the tool and data, and writing up. The core fits comfortably; the extension is planned, not improvised.

How much will this cost Lossfunk?

Very little. Hosting is free. Participants' own code runs in their browser, and grading happens offline on my machine, so there's no server running anyone's code. Model answers and AI checkers are expected to cost a few dollars, capped at US$50. The real cost is the panel members' time: about 3 hours each.

Is writing 32 tasks realistic?

About 2–3 weeks, including hidden tests and second-person checks, in month 1.

What's the biggest risk?

Not having enough people. I state the assumption openly and raise it early. The project produces the instrument, the task set and the AI-checker comparison regardless.

What will exist at the end?

An open harness, the task set with planted and natural errors and hidden tests (except the held-out set: the spares used for the realism check, whose task text stays private), anonymised raw data, the realism ratings with task IDs and error types, the AI-checker comparison, and a write-up.

Interpretation

What result would surprise you?

Checking taking as long as doing even on easy tasks; or subtle errors almost never getting through; or humans and AI checkers missing completely different errors.

If checking turns out cheaper than doing, is that the end of it?

No. Wrong answers that get through are a quality cost. Faster isn't a win if more wrong answers get through: Ipeirotis & Zheng found an AI tool made people faster without more correct answers per hour.

Can you say stronger models are harder to check?

Not as a causal claim. With a few errors per model, I can only describe what happened: on these tasks, a given model's wrong answers were caught more or less often.

How could someone else build on this?

Use the open harness with more people, other kinds of tasks or newer models. Because the protocol is declared in advance, their results can be compared with these.

Six months

Overview

The fellowship runs January to June 2027, my final semester, which is a full-time internship semester. I'd be on-site full-time for all six months.

The plan has four phases:

Months Phase What it produces
January Build the tasks 32 checked tasks, model answers, graded
February Build the instrument Working harness, planted errors, allocation, AI-checker results, pre-registration
March–April Run the panel Human data from 3–4 people
May–June Analyse, extend, release Results, one extension, open release, write-up

The core study fits comfortably in four months. The last two go into analysis, one planned extension, and releasing everything openly.

Assumptions

What success looks like

Level What exists at the end
Minimum The open harness, the checked task set and the AI-checker comparison, released, none of which depend on the panel; plus a pilot-sized human dataset if at least some people take part.
Target The full core study with 3–4 people (ideally 4): verification ratios per person, false-accept rates, H1–H3 tested descriptively, RQ3 and RQ4 described, released with a write-up.
Stretch The target, plus one of the May extensions, plus a workshop paper or arXiv preprint.

The minimum is reachable even if recruitment fails, because the instrument, the tasks and the AI-checker comparison don't depend on the panel.

Milestones

Month Goal Deliverable
January Tasks written, checked and answered by models 32 tasks (24 used, 6 spare, 2 practice) with specs, worked examples and hidden tests; model answers graded
February Instrument ready Harness, planted errors, pairs and allocation, ethics approval, OSF pre-registration, then AI-checker results v1; panel provisionally confirmed
March Panel, first half Screening and the walkthrough person's realism ratings (week 9); practice session and sessions 1–2 for every person; weekly data checks
April Panel, second half Sessions 3–4 (and a 5th if needed), realism ratings, debriefs; complete dataset
May Analysis and one extension Full analysis as pre-registered; human vs. AI-checker comparison; extension started
June Release and write-up Open release of harness, all 24 used tasks, anonymised data and analysis code (the held-out set, the spares used for the realism check, keeps its task text private; its realism ratings are released with task IDs and error types); write-up; talk at Lossfunk

Month by month

January: build the tasks

February: build the instrument

March: panel, first half

April: panel, second half

May: analyse and extend

June: release and write up

Change direction if…

When Signal What I'd do
January, week 1 Lossfunk can't confirm any panel members Treat the instrument, the tasks and the AI-checker comparison as the main results; look for volunteers from residency cohorts
End of January Tasks are taking much longer than planned Write 24 + 4 spares instead of 6; never cut the second-person check. With only 4 spares, any drop may push the realism check onto tasks people did
End of February Ethics approval not granted yet Keep building, and run the AI checkers. If it isn't granted by mid-March, the human part shrinks to a pilot
February, walkthrough Tasks take far longer or shorter than 5–10 minutes Rewrite them using the spares before any panel session
February, before the panel Walkthrough timings or model error rates suggest tasks are too easy Rewrite affected tasks using the spares before any panel session
Mid-March Sessions often run past 60 minutes Move to 2 pairs per session and add sessions. This is a contingency declared in the pre-registration; the pairs and allocation don't change
End of April Fewer than 3 people completed all their sessions The human part becomes a pilot; the instrument, the tasks and the AI-checker comparison become the main results
January–April Strong models make almost no natural errors Use planted errors for the main questions, and report RQ3 as a limitation. Never relabel planted errors as a model comparison
April Planted errors rated at least 1 point less realistic Report them separately and base conclusions mainly on natural errors

Risks

Risk Likelihood What I'd do
Not enough people for the panel High Confirm in week 1; screen 5 to get 3–4; the minimum success level doesn't depend on the panel
Someone drops out partway Medium Short, flexible sessions; a reserve from screening only before their first session; report what was completed
Ethics approval is slow Medium Apply in week 1; build everything else meanwhile
Writing 32 good tasks takes longer Medium About 11 a week; reduce spares if needed, never the checks
Harness bugs or timing errors Medium Walkthrough first; keep raw event logs; weekly data checks
Hidden tests leak Low Never sent to the browser; grading only offline
Models make too few natural errors Medium Planted errors cover the main questions; RQ3 becomes descriptive or a limitation
Planted errors look fake Medium Made from real model answers; realism ratings; report separately if needed
People improve with practice High Expected, tracked session by session, and reported as a finding
A model is retired or changes Low Exact versions recorded; all answers saved in January
Results are small and uncertain High Stated plainly as pilot-sized; the instrument is the lasting output
Released tasks end up in future training data Medium Keep the held-out set private: the spares used for the realism check

What I need from Lossfunk

  1. People: 3–4 community members, about 3 hours each, in March and April (about 5 hours each if possible).
  2. Ethics: advice on which review process to use.
  3. A small budget for API calls: a few dollars expected, capped at US$50, for model answers and AI checkers.
  4. A mentor: a short check-in every week or two.
  5. A quiet room for sessions.
  6. Optional: a small voucher budget, or access to residency cohorts, to widen the panel.

Budget

Item Cost
Participants Volunteers, no fee
Hosting the harness Free tier
Running participants' code Free: it runs in their own browser
Model answers and AI checkers A few dollars expected, capped at US$50
Grading Free: offline on my machine

How I'll work

Before January (only if accepted)

I'll be finishing my seventh semester, so this is light and optional: - read more deeply on the key papers; - draft a few practice tasks; - find out which ethics review applies, so I can apply in week 1; - find out what PES needs for the internship to count (offer letter, any reports).

Pilot

Status

No pilot has been run yet. The design requires ethics approval before any data is collected from people, and the panel comes from the Lossfunk community.

Planned

A software-only walkthrough of the harness (February, week 7), then screening and realism ratings after ethics approval (March, week 9). See PLAN.md.

Already tested

The spot-the-bug puzzle on this page, whose bug was verified with tests.

Working with AI

Caught by me

Caught by checking the AI's work