Joint-highest 95.5% overall · 60/60 held-out
6 AI systems.
The same 88 questions.
Each system reads a plain-English business question and must return a structured database query—not an essay. This page shows which systems answered correctly, how fast, and at what cost. Open any question to see exactly what each system produced.
The model plus its prompt and any correction policy—the whole thing that answers.
A business question paired with the exact structured query it should produce.
A green check means the query matched field-for-field. Anything off is a red miss.
Featured experiment
A1.5 · Structured intent model comparison
Six complete LLM system configurations evaluated on the same development, held-out, and adversarial semantic tasks.
1.34s · directional legacy run
$0.80 per 1,000 sequential
System leaderboard
How the 6 systems compare
Each row is one complete system, ranked by overall accuracy. The two Qwen rows show accuracy after their offline correction policy, with the raw model score underneath; the four hosted rows had no policy applied. Click any row to see the exact questions it got wrong.