New here? 50% off your first month with code WELCOME50
repryntt

autonomy, measured

We run our own AI workforce. Here's its flight recorder.

This is real telemetry from our own always-on instance of the open-source repryntt engine โ€” an AI employee that has been operating continuously since March: researching, writing, running commands, reading email, and keeping its own memory. Every action is logged. These are the logs.

March 14 โ€“ July 7, 2026 ยท regenerated July 7, 2026

198,008

logged actions

every step, on the record

101

days of operation

not a demo run

119,932

tool calls

real work, not chat

16,806

model calls

12 models ยท 6 providers

the work

What an AI employee actually does all day.

Not answers โ€” actions. The distribution below is the working life of one autonomous agent, straight from its tool-call log.

Journaling & memory writes11,105

it writes its own working memory more than it does anything else

Reading files8,592
Writing files7,354
Web research13,516

searches + page reads across two search hands

Scraping & reading pages3,891
Terminal commands3,742

real shell, real machine

Market & chain queries4,514

on-chain and market data feeds

Reading email1,944

its own inbox, triaged daily

the brains

Every model we've run it on.

repryntt is bring-your-own-key โ€” so we've run the same workforce on a dozen models across six providers. The engine doesn't care whose model is underneath. That's the point.

Mistral Small 4 (119B)7,985
Mistral Large 3 (675B)2,448
NVIDIA Nemotron 49B (v1.5 + v1)3,332
Grok 4.1 Fast / 4.31,335
Gemini 2.0 / 2.5 Flash944
Claude Opus 4.6471
NVIDIA Nemotron 3 (120B)155
Llama 3.3 70B117
Claude Fable 511

next

Benchmarks for autonomous work.

Standard AI benchmarks measure how well a model answers questions. Nobody publishes how well models complete real work end to end โ€” tasks with a verifiable outcome: the file written, the test passed, the email sent, the build green.

As of July 2026 our engine records the outcome of every action โ€” success, latency, and token cost, per model, per task type. As that data accumulates across models, this page will grow the numbers we most want to see ourselves: which models finish the highest share of real tasks, at what speed, at what cost. Same workforce, same tasks, different brains โ€” measured, not marketed.

coming

Task completion rate

Share of real tasks each model completes with a verified outcome โ€” the number that actually predicts an autonomous workforce's usefulness.

coming

Speed per task

Median latency from "task picked" to "outcome verified," per model โ€” how long your workforce makes you wait.

coming

Cost per completed task

Tokens spent per verified outcome โ€” because BYOK means this number lands on your bill, not ours.