SOLO BUILD BY ARIEL MAGALSO · LIVE PIPELINE — NOT A SCRIPT

Verdict.

//EVALS

Reliability evidence

A scorecard, published as-is. 60 cases test the real pipeline — real Anthropic calls, real Postgres writes — across all seven categories in the project spec, at the spec's own target counts. 18 were tuned directly against while building the rule engine; the other 42 were designed from the scoring spec and run once, held out from tuning. Both numbers are shown below, not blended into one.

95%
Accuracy across 60 cases
0
False-score rate — a number emitted on evidence that should have refused
2
False-refusal rate — a refusal on evidence that should have scored
1.2s
Mean latency per case · ~$0.60 total

//DEV SET VS. HELD-OUT SET

100%
Dev set (tuned against) · 21/21 cases
92%
Held-out set (not tuned against) · 36/39 cases

//BY CATEGORY

Category Cases Passed Accuracy
Adversarial (prompt injection) 5 5 100%
Disqualified 10 10 100%
Duplicate / merge review 5 5 100%
Insufficient evidence 5 5 100%
Needs review 10 10 100%
Nurture 10 8 80%
Sales-ready 15 14 93%

//FAILURES — SHOWN, NOT HIDDEN

eval-sr-14 (Sales-ready)

band=needs_review expected=sales_ready

eval-nu-08 (Nurture)

outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture

eval-nu-09 (Nurture)

outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture

Last run: 2026-08-15 18:59 UTC · eval set v1-60cases-python · model: claude-haiku-4-5-20251001