RQ1
Time
How does check time compare with do time, and does it change with difficulty?
- H1 Checking is faster on average, but the ratio rises with difficulty.
Exhibit A · AI-written
def second_largest(nums): if len(nums) < 2: return None first = second = None for n in nums: if first is None or n > first: second = first first = n elif second is None or n > second: ✗ bug second = n return second
Is this function correct?
nums, or None if there is none.second_largest([3, 1, 4, 2]) returns 3.00:00 time spent checking
The function is incorrect. The bug is on the elif line. A repeated largest value is kept as the second-largest: [5, 5, 3] returns 5, not 3.
Answer: incorrect. The bug is on the elif line. A repeated largest value is kept as the second-largest: [5, 5, 3] returns 5, not 3.
You just paid a verification cost:00:00. This project measures it.
When you could do a task yourself, does checking an AI's attempt cost less time than doing it? And how often does a wrong answer get through?
§1
Delegating to AI assumes that checking its output is cheaper than doing the work.
Illustration: these are inputs you choose, not study results.
Delegating saves 2 min
Delegating saves time only if check time + (non-accept rate × do time) < do time.
Non-accept: a reject, a timeout, or "unverifiable": each means redoing the task.
§2
Other fields have measured this. In code, the studies I found measure only part of it.
| Translation & radiology | Checking only | With vs. without AI | This study | |
|---|---|---|---|---|
| Same people do and check | Yes | No | Partly | Yes |
| Exact answers (automatic checking) | No | Yes | Partly | Yes |
| Verdict only (no rewriting) | No | Yes | No | Yes |
| False accepts measured | Partly | Yes | Partly | Yes |
| Code tasks | No | Partly | Yes | Yes |
| Key papers |
|
|
|
Not found in this search: a code study where the same people solve tasks and judge AI answers to matched tasks, with answers checked automatically.
§3
Each person solves some tasks and checks AI answers to matched ones, so each person is their own baseline.
Screening
Practice
4 sessions
Realism ratings
Debrief
3–4
people in the panel
24
tasks per person (12 do, 12 check)
10min
check budget
50%
of answers wrong
2–3
models write the answers
~3h
per person
flowchart TD A["Write 32 code tasks (24 used, 6 spare, 2 practice)"] --> B[Model answers: 1 correct and 1 wrong per task] B --> C[Pair tasks] C --> D[Do: solve one task of each pair] C --> E[Check: judge the AI answer to its partner] P["Panel: 3–4 people, sessions of 3 pairs (max 60 min)"] --> D P --> E D --> G[Do time and correctness] E --> H[Check time, verdict and confidence] G --> I[Per-person ratios and false accepts] H --> I B --> J["AI checkers: 3–5 models judge every answer 5 times"] H --> K[Errors missed by humans vs. AI checkers] J --> K
check time + (non-accept rate × do time) < do time.§4
Four questions: time, errors that get through, differences across models, and AI checkers.
RQ1
How does check time compare with do time, and does it change with difficulty?
RQ2
How often are wrong AI answers accepted within budget?
RQ3
Are a stronger model's own errors caught less often, or more slowly?
RQ4
Do AI checkers miss the same errors human checkers miss?
§5
The core study fits in four months. The last two go into analysis, one extension and an open release.
January
Build the tasks
32 checked tasks, model answers, graded
February
Build the instrument
Harness, planted errors, allocation, pre-registration, AI-checker results
March–April
Run the panel
Human data from 3–4 people
May–June
Analyse, extend, release
Results, one extension, open release, write-up
CP0 January, week 1
Confirm who can join the panel, and when.
CP1 End of February
Ethics approved, pre-registration done, harness working, at least 3 people provisionally confirmed.
CP2 Mid-March
Do sessions fit in 60 minutes? If not, use the pre-registered 2-pair contingency.
CP3 End of April
How many people completed all their sessions?
Deliverable 32 tasks (24 used, 6 spare, 2 practice) with specs, worked examples and hidden tests; model answers graded.
Deliverable harness, planted errors, pairs and allocation, ethics approval, OSF pre-registration, then AI-checker results.
Deliverable practice session and sessions 1–2 for every person.
Deliverable complete dataset.
Deliverable full analysis; extension started.
Deliverable open release and write-up.
Minimum
The open harness, the checked task set and the AI-checker comparison, released. Plus a pilot-sized dataset if some people take part.
Target
The full core study with 3–4 people (ideally 4), released with a write-up.
Stretch
The target, plus one extension, plus a workshop paper or arXiv preprint.
The minimum doesn't depend on the panel.
§6
The biggest risk is not having enough people. The minimum result doesn't depend on them.
| Risk | Likelihood | What I'd do |
|---|---|---|
| Not enough people for the panel | High | Confirm in week 1; screen 5 to get 3–4. |
| Results are small and uncertain | High | Stated plainly as pilot-sized. |
| Ethics approval is slow | Medium | Apply in week 1; build meanwhile. |
| Models make too few natural errors | Medium | Planted errors cover the main questions. |
| Planted errors look fake | Medium | Realism ratings; reported separately if needed. |
If Fewer than 3 people complete all their sessions
Then The human part becomes a pilot.
If Ethics isn't approved by mid-March
Then The human part shrinks to a pilot.
If Sessions often run past 60 minutes
Then Move to 2 pairs per session, as pre-registered.
If Tasks look too easy before the panel
Then Rewrite them using the spares.
§7
Mainly people's time: 3–4 community members, about 3 hours each, in March and April.
§8
Everything is released openly, so anyone can run the study with more people.
Held out The spares used for the realism check stay private; their ratings are released.
§9
The questions a reviewer is most likely to ask, answered briefly.
Partly. There, people rewrite the output and humans judge quality. Here, people give a verdict only, and hidden tests check answers.
Checking time alone doesn't show whether delegating helped. Five minutes to check is good if doing takes thirty, bad if it takes four.
For this panel, yes. Each person does 24 tasks and counts as a replication. Error rates are pilot-sized, and I say so.
They might be. Each edits a correct model answer; if it targets an edge case, the spec states that case. If rated at least 1 point less realistic, they're reported separately.
It might. Checkers aren't told the rate. I report the false-accept rate per wrong answer, which depends less on the mix.
Not as a causal claim. With a few errors per model, I can only describe what happened on these tasks.
§10
I learn what I need, I build the measurement, and I scope to what's real.
Before taking any ML courses, I taught myself enough to build an unsupervised anomaly-detection pipeline for ship-tracking data end to end, and published it as first author (DoSCI-2026).
For my Contract Risk Analyzer, which I designed, built and deployed myself, I wrote my own 90-clause benchmark to test whether it actually worked.
This proposal was rebuilt twice: once after a literature review narrowed the claim, and once to fit the 3–4 people I can actually count on. Every figure from prior work was checked against the original papers.
§11
I used AI to draft and review. I checked its work against the papers and pushed back where it was wrong.