//EVALS
Reliability evidence
A scorecard, published as-is. 60 cases test the real pipeline — real Anthropic calls, real Postgres writes — across all seven categories in the project spec, at the spec's own target counts. 18 were tuned directly against while building the rule engine; the other 42 were designed from the scoring spec and run once, held out from tuning. Both numbers are shown below, not blended into one.
//DEV SET VS. HELD-OUT SET
//BY CATEGORY
| Category | Cases | Passed | Accuracy |
|---|---|---|---|
| Adversarial (prompt injection) | 5 | 5 | 100% |
| Disqualified | 10 | 10 | 100% |
| Duplicate / merge review | 5 | 5 | 100% |
| Insufficient evidence | 5 | 5 | 100% |
| Needs review | 10 | 10 | 100% |
| Nurture | 10 | 8 | 80% |
| Sales-ready | 15 | 14 | 93% |
//FAILURES — SHOWN, NOT HIDDEN
band=needs_review expected=sales_ready
outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture
outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture
Last run: 2026-08-15 18:59 UTC · eval set v1-60cases-python · model: claude-haiku-4-5-20251001