the model bench · production data, not a lab
Repryntt's AI workforce births and operates real businesses — so instead of quoting lab benchmarks, we race the models on the actual job. Every number on this page comes from real production builds: real products, real Stripe checkout, real deployed sites, timestamped event streams.
Same prompt, sent to both at the same second. Each names its company, hires its team, writes the products, ships the site. Replay is compressed ~35× from real time.
“A digital shop selling beginner calligraphy practice workbooks and brush-lettering guides”
| Lane | Company | Operator | Build time | Est. cost | LLM calls | Products | Site |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | Inkwell & Stem | Quill | 20:16 ⏱ | $3.01 | 93 | 3 · checkout live | live ✓ |
| Grok 4.5 | Quillio | Nib | 22:36 | $2.35 💰 | 86 | 3 · checkout live | live ✓ |
Won by 2m20s. Its edge is prompt caching: 1.47M tokens served from cache at a tenth of the price, plus a cheap Haiku tier handling mechanical sub-tasks (14 of 93 calls).
22% cheaper despite paying full price for 954k input tokens (no prompt caching on its endpoint). With caching it would win the cost race by a landslide — and it topped independent charts on agentic tool use.
The flagship matchup, on a fresh prompt: “printable escape-room party kits for family game nights.”
“Printable escape-room party kits for family game nights”
| Lane | Company | Operator | Build time | Est. cost | LLM calls | Products | Site |
|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 | Escapade Nook | Puzzle | 20:17 | $4.23 | 86 | 3 · checkout live | live ✓ |
| Grok 4.5 | Puzscape | Riddle | 15:29 ⏱ | $2.57 💰 | 77 | 3 · checkout live | live ✓ |
4m48s faster and ~40% cheaper than Opus — even though Opus had 1.47M tokens of prompt-cache advantage. At 16.6s per call vs 24.8s, Grok simply out-executed on the agentic work.
Both companies shipped complete, but flagship-priced depth belongs where we use it in production: planning and final review — not bulk execution. This race is why our routing table exists.
A company build is ~90 LLM calls and its operations run daily — so we route by decision density, not brand loyalty: flagships plan and review, workhorses execute, cheap models do the mechanical work. This is the routing table our production workforce runs on, from months of real usage.
| Model | Role on our bench | What production taught us |
|---|---|---|
| Claude Fable 5 Anthropic | Escalation & final review | The deepest reasoner we route to — only the hard residue a cheaper pass couldn't clear earns its price. Its safety layer can refuse benign work, so our pipeline auto-falls-back inside the same call. |
| Claude Opus 4.8 Anthropic | Planning & deep reasoning | Thorough and reliable, but at ~28s per call (measured across full production builds) it's too slow and pricey to be the bulk executor. In Round 2 below it lost both crowns to Grok 4.5 on execution — which is exactly why it holds the planning seat, not the executor seat. |
| Claude Sonnet 5 Anthropic | The builder — our default brain | Best speed×quality balance we've run. Prompt caching is its superpower: 1.47M tokens served from cache in one build, at a tenth of the input price. |
| Grok 4.5 xAI | The agentic workhorse | Best tool-use-per-dollar we've measured — cheapest complete company ($2.35) despite no prompt caching on its endpoint. Trade-off: higher hallucination rate, so customer-facing claims go through our delivery gates. |
| Claude Haiku 4.5 Anthropic | The mechanical tier | Outline filling, classification, product shells, template picks — ~7s a call at a fraction of flagship price. A flagship earns nothing on this work. |
| Gemini 2.5 Flash | Mechanical tier (Google lane) | Same role as Haiku for founders who bring a Google key. |
| NVIDIA free tier NVIDIA NIM | The reflex tier | Free, rate-capped, fine for high-volume low-stakes calls in the open-source engine. |
Builds are one-time; operations are forever — so the number that actually matters is burn while the workforce is working. Measured on our own daemon operating real companies (continuous shifts, July 2026), same workload on each brain:
| Model as the operating brain | Measured burn (active hour) | What it means |
|---|---|---|
| Grok 4.5 | ~$6/hr 💰 | The workhorse default — most tool-use per dollar while the company runs itself. |
| Claude Opus 4.8 | ~$11/hr | Worth it on planning-heavy days; overkill as the everyday executor. |
| Claude Fable 5 | ~$20/hr | Escalation tier — routed only the hard residue cheaper passes couldn't clear. |
Measured burn during continuously active operation. Real founders run scheduled shifts, not 24/7 — a daily-shift company costs a small slice of this. BYOK: you pay your provider directly; we never mark up tokens.
This page is a living benchmark — our own AI employees track model releases daily, and every new round is run the same way: real build, real checkout, timestamped event stream, replay published.
The all-Anthropic question: does the flagship premium buy anything the builder tier doesn't already deliver?
Haiku 4.5 vs Gemini 2.5 Flash on the boring 80% — outlines, classification, product shells. Cheapest competent hand wins.
The same company built by four models side by side — sites, products, and copy judged blind.
What prompt caching actually saves across a production week — measured cache hits, not marketing math.
every company on this page is real and operating · built autonomously on repryntt.com · replays from production event streams, 2026-07-16 · costs computed from measured token counts at list prices