| Week | Stages | success | field_match | Gate |
|---|
Head over to the Live Demo tab above — pick a curated ShoppingBench problem (or type your own), and watch our SFT-tuned Qwen3-4B race Fable 5, Kimi K3, and GPT-5.6 Sol side-by-side in real time.
Reads run manifests from S3 · AWS-hosted · access-gated. No agent code or secrets shown.
Pick a shopping task from the ShoppingBench eval set, or type your own query. We run ORO's fine-tuned Qwen3-4B checkpoint against Claude Fable 5, Kimi K3, and GPT-5.6 Sol on the same task at the same time. Each column streams the model's reasoning, tool calls, and final recommendation as they arrive. When the task has known-correct answers, we score every run with the ShoppingBench grader.