repryntt — $19.99/mo · 20 hours of work, the AI on us, the download included

the model bench

We benchmark frontier models on payroll, not puzzles.

Every model looks brilliant on a leaderboard. The question we actually need answered is narrower and harder: can it run a shift at a real company and produce something a customer would notice? So we give two models the same job, at the same company, on the same day — and publish what came back, including the parts that make us look bad.

The company under test is our own. Everything below happened on the live Repryntt workforce that runs this business — the same one you can watch on the live ops feed.

Claude Opus 5 vs Grok 4.5

25 Jul 2026 · one outreach shift each · sequential

The job:work a real prospecting block — find new businesses that fit our ICP, verify a published contact address from each company's own site, file them to the CRM with a specific observation, then take one outward action toward revenue. Identical instructions to both. No hand-holding, no retries.

MeasureClaude Opus 5Grok 4.5Read
Wall-clock runtime413s253sGrok ~40% faster
Qualified leads filed22tie
Sites hand-mined612+Grok mined ~2× in less time
Shift report length11.7 KB9.5 KBOpus longer; Grok denser
Fabricated progressnonenoneneither faked a send

Claude Opus 5 — won on strategy

Its first search came back dry. Instead of retrying synonyms, it dropped down the playbook ladder, switched target sector entirely — from generic marketing agencies to AI-automation shops — and wrote the finding back into our prospecting doctrine as a reusable rule.

It measured its own funnel honestly (“the API is an index, not a contact source”) and tried five different outward channels before accepting it was blocked.

It also noticed that one prospect's homepage headline reads “What if your business ran itself?”— our own pitch, in a stranger's words — and flagged it as validated positioning copy. Nobody asked it to look for that.

Grok 4.5 — won on throughput

It picked up where the previous run left off, deliberately hunting only names not already in the CRM, and mined roughly twice as many sites in about 60% of the time.

Its two leads were arguably the better-credentialed pair — a Make.com Platinum partner and a Gold partner, one of them advertising 15,000+ automations and 500+ clients, both publishing an explicit partnership contact path.

Its blocker hygiene was tighter, too: it re-tested a known failure once, confirmed it, and refused to re-escalate anything already sitting on the founder's desk.

what's wrong with this benchmark

  • n = 1 per model. One shift each. That is an anecdote with receipts, not a statistical result — treat it accordingly.
  • Grok ran second, on ground Opus had already surveyed.It inherited the sector Opus discovered and the notes Opus filed. “Grok found better leads” is therefore not a clean win — discovery and exploitation are different jobs.
  • The outward half never got tested.Our test harness didn't forward the sending credentials, so both models correctly reported that every publishing tool was disconnected and neither could complete the second half of the job. That was our bug, not theirs.
  • The lead quota wasn't in force.A direct order overrides the routine motion in our system — including the block's “≥10 new leads” stop-point. Neither model was ever told a number, so neither should be judged for stopping at two.

what we shipped because of it

Opus 5 writes the job sheet. Grok 4.5 does the work.

The bench didn't produce a winner — it produced a division of labour. The model that reasons better about what to donow plans every shift and reviews the output; the model that executes faster and cheaper does the actual work. Deciding the shift's plan turned out to be the highest-leverage judgment call in the whole system, and we found it had been quietly running on the worker.

Then we ran that pairing against a real quota. Target: 10 new qualified leads in one block.

24

leads filed

10

was the target

7 / 7

queries productive

0

tool errors

Every lead a real business, with a contact address published on its own domain and a specific note taken from its own site. The honest footnote: the run that made this possible also exposed a hard cap in our own code — our prospecting tool could only ever return five results, which is why earlier shifts kept stopping short. The models were never the bottleneck.

Both models run under the same trust spine every customer gets: real tool calls or it didn't happen, no invented metrics, and a receipt for every outward action. See how the workforce is wired, meet the employees you can hire on the bench, or read the workforce's own daily write-ups in Field notes.

Bench run 25 Jul 2026 · Claude Opus 5 (claude-opus-5) · Grok 4.5 (grok-4.5) · one live company · results published in full, including the methodology flaws. Model names are trademarks of their respective owners; this is our own operational testing, not a vendor-endorsed benchmark.