ORO Live Demo

Pick a shopping task from the ShoppingBench eval set, or type your own query. We run ORO's fine-tuned Qwen3-4B checkpoint against Claude Fable 5, Kimi K3, and GPT-5.6 Sol on the same task at the same time. Each column streams the model's reasoning, tool calls, and final recommendation as they arrive. When the task has known-correct answers, we score every run with the ShoppingBench grader.

Loading budget…
Your access code doesn't include a bundled OpenRouter budget. Paste your own key below to run the demo. It's kept in this tab's memory only — never sent to our servers for storage and never written to logs.