footnote1 running answers, grounded in research

Evals

Judge

LLM judge vs 20 hand labels.
configurationTPRTNRfalse alarms
v0 rubric, bare0.93 (13/14)0.00 (0/6)0/14
v0 rubric + reasoning0.930.33 (2/6)0/14
v1 rubric, bare0.930.33 (2/6)0/14
v1 rubric + reasoning0.930.67 (4/6)1/14

n=6 negatives. Wilson 95% on 4/6: 0.30–0.90.

Retrieval

Retrieval recall. n=15 questions, 7 skipped.
armrecall@5recall@10
lexical (ts_rank)54.2%88.1%
vector (bge-small)56.0%84.2%
hybrid (RRF, k=60)71.8%84.2%

Relevance labels are machine-generated, pending hand relabelling. Vector and hybrid are likely flattered.

Triage

Triage vs hand labels. Stratified: 60 accepted, 60 rejected.
stratumquestion it answersstatus
rejected papersshare wrongly thrown away — the silent errorlabelling, 6/60
accepted papersshare wrongly kept — the noisy errorlabelling, 6/60

Started after triage rejected 3 of the 5 hand-picked seed papers.

as of 2026-08-23method