Loading…
Recorded runs of the advisor (v3.1) over Sanity Context MCP on project onwa0wvs, dataset v2.
No model runs on this page: every tool call, every validator round and every answer is exactly what was recorded.
All 27 runs of the measured run (9 scenarios × 3, blind grader total 59/78) — nothing picked or hidden, including weak answers and answers that came out in Russian (a language bug, fixed later by a code check). Long tool results are cut. Code & method