PURE · AI Model Evals
AI Cost Admin

PURE · AI · Quality-per-tier

AI model evals

"Best results, not just cheapest" — made permanent. Each task type is scored at Haiku, Sonnet and Opus against the live PURE canon, so routing is proven by numbers: bump a task up if the cheap tier underperforms, drop it down if the expensive tier isn't earning its cost. Smart routing saves Pure Dollars (PD) by using the cheapest tier that still passes.

sample eval scores — run the harness for live
Avg routed score
Flag: bump up
Flag: save (drop)
Task types
Nightly over poppy_evals; scores 0–100 per tier. Routing recommendation is auto-derived.
Task typeHaikuSonnetOpusRoutedRecommendation