Exhibit A · AI-written

second_largest.pypython
def second_largest(nums):    if len(nums) < 2:        return None    first = second = None    for n in nums:        if first is None or n > first:            second = first            first = n        elif second is None or n > second: ✗ bug            second = n    return second

Is this function correct?

Spec
return the second-largest distinct value in nums, or None if there is none.
Example
second_largest([3, 1, 4, 2]) returns 3.

00:00 time spent checking

Measuring the Human Cost of Verifying AI Output

When you could do a task yourself, does checking an AI's attempt cost less time than doing it? And how often does a wrong answer get through?

§1

Why it matters

Delegating to AI assumes that checking its output is cheaper than doing the work.

  • If checking takes about as long as doing, or wrong answers slip through, delegation saves little or nothing.
  • Developers were 19% slower with AI while believing they were faster (METR 2025).
  • AI code reviews saved no time (Tufano et al. 2025).

Illustration: these are inputs you choose, not study results.

8 min
4 min
25%
Doing it yourself 8 min
Checking + redoing 6 min

checking redoing after a non-accept

Delegating saves 2 min

Delegating saves time only if check time + (non-accept rate × do time) < do time.

Non-accept: a reject, a timeout, or "unverifiable": each means redoing the task.

§2

What's known

Other fields have measured this. In code, the studies I found measure only part of it.

Translation & radiologyChecking onlyWith vs. without AIThis study
Same people do and checkYesNoPartlyYes
Exact answers (automatic checking)NoYesPartlyYes
Verdict only (no rewriting)NoYesNoYes
False accepts measuredPartlyYesPartlyYes
Code tasksNoPartlyYesYes
Key papers
  • Green et al. 2013
  • Sarti et al. 2022
  • Acosta et al. 2024
  • Kaufman et al. 2026
  • Wen et al. 2024
  • Becker et al. (METR) 2025
  • Tufano et al. 2025
  • Ipeirotis & Zheng 2025

Not found in this search: a code study where the same people solve tasks and judge AI answers to matched tasks, with answers checked automatically.

§3

The study

Each person solves some tasks and checks AI answers to matched ones, so each person is their own baseline.

One person's journey

  1. Screening

    • a 10-minute test, done online beforehand
  2. Practice

    • 2 practice tasks, not used in the study
  3. 4 sessions

    • 3 pairs each: solve 3, check 3
    • about 45 min, hard stop at 60
  4. Realism ratings

    • about 8 answers to spare tasks
  5. Debrief

    • explains the planted errors
  • Nobody solves and checks the same task.
  • No right/wrong feedback during the study.
  • 10-minute check budget; 20-minute do cap.

3–4

people in the panel

24

tasks per person (12 do, 12 check)

10min

check budget

50%

of answers wrong

2–3

models write the answers

~3h

per person

How it fits together

flowchart TD
  A["Write 32 code tasks (24 used, 6 spare, 2 practice)"] --> B[Model answers: 1 correct and 1 wrong per task]
  B --> C[Pair tasks]
  C --> D[Do: solve one task of each pair]
  C --> E[Check: judge the AI answer to its partner]
  P["Panel: 3–4 people, sessions of 3 pairs (max 60 min)"] --> D
  P --> E
  D --> G[Do time and correctness]
  E --> H[Check time, verdict and confidence]
  G --> I[Per-person ratios and false accepts]
  H --> I
  B --> J["AI checkers: 3–5 models judge every answer 5 times"]
  H --> K[Errors missed by humans vs. AI checkers]
  J --> K

Key terms

Verification ratio
Check time ÷ do time, on matched tasks. Below 1 means checking is faster.
False accept
A wrong answer judged correct.
Premature closure
A wrong verdict given at 100% confidence.
Break-even
Delegating saves time only if check time + (non-accept rate × do time) < do time.

§4

Questions and hypotheses

Four questions: time, errors that get through, differences across models, and AI checkers.

RQ1

Time

How does check time compare with do time, and does it change with difficulty?

  • H1 Checking is faster on average, but the ratio rises with difficulty.

RQ2

Errors that get through

How often are wrong AI answers accepted within budget?

  • H2 Subtle errors are accepted much more often than obvious ones.
  • H3 Wrong answers are accepted faster than correct ones are rejected.

RQ3

Across models (exploratory)

Are a stronger model's own errors caught less often, or more slowly?

  • H4 Fewer errors, but each more likely to get through. Uncertain.

RQ4

AI checkers (supporting)

Do AI checkers miss the same errors human checkers miss?

  • — Open question, no prediction.

§5

Six-month plan

The core study fits in four months. The last two go into analysis, one extension and an open release.

  1. January

    Build the tasks

    32 checked tasks, model answers, graded

  2. February

    Build the instrument

    Harness, planted errors, allocation, pre-registration, AI-checker results

  3. March–April

    Run the panel

    Human data from 3–4 people

  4. May–June

    Analyse, extend, release

    Results, one extension, open release, write-up

  1. CP0 January, week 1

    Confirm who can join the panel, and when.

  2. CP1 End of February

    Ethics approved, pre-registration done, harness working, at least 3 people provisionally confirmed.

  3. CP2 Mid-March

    Do sessions fit in 60 minutes? If not, use the pre-registered 2-pair contingency.

  4. CP3 End of April

    How many people completed all their sessions?

January
  • Week 1: confirm the panel, apply for ethics, choose the models.
  • Weeks 2–3: write the 32 tasks, about 11 a week.
  • Week 4: a second person checks every task; each model answers each task 5 times, graded offline.

Deliverable 32 tasks (24 used, 6 spare, 2 practice) with specs, worked examples and hidden tests; model answers graded.

February
  • Week 5: choose wrong answers, plant errors, pair tasks, build the allocation.
  • Week 6: build the harness and the offline grading.
  • Week 7: walkthrough to test the software only.
  • Week 8: pre-register, then run the AI checkers.

Deliverable harness, planted errors, pairs and allocation, ethics approval, OSF pre-registration, then AI-checker results.

March
  • Week 9, after ethics approval: screening, and the walkthrough person's realism ratings.
  • From week 10: practice session, then sessions 1 and 2.
  • Weekly checks on data quality only, not results.

Deliverable practice session and sessions 1–2 for every person.

April
  • Sessions 3 and 4, and a 5th if needed.
  • Realism ratings on spare tasks, then a debrief.
  • Grade all submitted answers offline.

Deliverable complete dataset.

May
  • Run the pre-registered analysis, H1 to RQ4.
  • Compare human and AI-checker misses.
  • Start one extension, chosen at checkpoint 3.

Deliverable full analysis; extension started.

June
  • Release the harness, the 24 used tasks, anonymised data and analysis code.
  • Write up the results honestly at their real size.
  • Give a talk at Lossfunk.

Deliverable open release and write-up.

What success looks like

  1. Minimum

    The open harness, the checked task set and the AI-checker comparison, released. Plus a pilot-sized dataset if some people take part.

  2. Target

    The full core study with 3–4 people (ideally 4), released with a write-up.

  3. Stretch

    The target, plus one extension, plus a workshop paper or arXiv preprint.

The minimum doesn't depend on the panel.

§6

Risks and what would change my mind

The biggest risk is not having enough people. The minimum result doesn't depend on them.

Top risks

RiskLikelihoodWhat I'd do
Not enough people for the panelHighConfirm in week 1; screen 5 to get 3–4.
Results are small and uncertainHighStated plainly as pilot-sized.
Ethics approval is slowMediumApply in week 1; build meanwhile.
Models make too few natural errorsMediumPlanted errors cover the main questions.
Planted errors look fakeMediumRealism ratings; reported separately if needed.

Change direction if

  • If Fewer than 3 people complete all their sessions

    Then The human part becomes a pilot.

  • If Ethics isn't approved by mid-March

    Then The human part shrinks to a pilot.

  • If Sessions often run past 60 minutes

    Then Move to 2 pairs per session, as pre-registered.

  • If Tasks look too easy before the panel

    Then Rewrite them using the spares.

§7

What I need from Lossfunk

Mainly people's time: 3–4 community members, about 3 hours each, in March and April.

  • People 3–4 members, about 3 hours each (5 if possible).
  • Ethics Advice on which review process to use.
  • API budget A few dollars expected, capped at US$50.
  • A mentor A check-in every week or two.
  • A quiet room For sessions.
  • Optional Vouchers or residency cohorts, to widen the panel.

§8

What will exist at the end

Everything is released openly, so anyone can run the study with more people.

  1. An open harness.
  2. All 24 used tasks, with hidden tests.
  3. Anonymised data and analysis code.
  4. The AI-checker comparison.
  5. A write-up, a talk and a short guide.

Held out The spares used for the realism check stay private; their ratings are released.

§9

Hard questions

The questions a reviewer is most likely to ask, answered briefly.

Q1 Isn't this already known from translation research?

Partly. There, people rewrite the output and humans judge quality. Here, people give a verdict only, and hidden tests check answers.

Q2 Why compare checking with doing?

Checking time alone doesn't show whether delegating helped. Five minutes to check is good if doing takes thirty, bad if it takes four.

Q3 With 3–4 people, is any of this meaningful?

For this panel, yes. Each person does 24 tasks and counts as a replication. Error rates are pilot-sized, and I say so.

Q4 Aren't planted errors unrealistic?

They might be. Each edits a correct model answer; if it targets an edge case, the spec states that case. If rated at least 1 point less realistic, they're reported separately.

Q5 Doesn't 50% wrong answers make checkers suspicious?

It might. Checkers aren't told the rate. I report the false-accept rate per wrong answer, which depends less on the mix.

Q6 Can you say stronger models are harder to check?

Not as a causal claim. With a few errors per model, I can only describe what happened on these tasks.

All 58 hard questions on the full page →

§10

Why I can run this

I learn what I need, I build the measurement, and I scope to what's real.

I learn what I need and ship it.

Before taking any ML courses, I taught myself enough to build an unsupervised anomaly-detection pipeline for ship-tracking data end to end, and published it as first author (DoSCI-2026).

I build the measurement, not just the system.

For my Contract Risk Analyzer, which I designed, built and deployed myself, I wrote my own 90-clause benchmark to test whether it actually worked.

I scope to what's real.

This proposal was rebuilt twice: once after a literature review narrowed the claim, and once to fit the 3–4 people I can actually count on. Every figure from prior work was checked against the original papers.

§11

How I worked with AI

I used AI to draft and review. I checked its work against the papers and pushed back where it was wrong.

  • I challenged the gap claim; a literature review narrowed it.
  • I rejected a design needing 84 sessions and rebuilt it for 3–4 people.
  • Browser grading would have leaked the hidden tests; grading is now offline.

Full log on the full page →