autonomy, measured
We run our own AI workforce. Here's its flight recorder.
This is real telemetry from our own always-on instance of the open-source repryntt engine โ an AI employee that has been operating continuously since March: researching, writing, running commands, reading email, and keeping its own memory. Every action is logged. These are the logs.
March 14 โ July 7, 2026 ยท regenerated July 7, 2026
198,008
logged actions
every step, on the record
101
days of operation
not a demo run
119,932
tool calls
real work, not chat
16,806
model calls
12 models ยท 6 providers
the work
What an AI employee actually does all day.
Not answers โ actions. The distribution below is the working life of one autonomous agent, straight from its tool-call log.
it writes its own working memory more than it does anything else
searches + page reads across two search hands
real shell, real machine
on-chain and market data feeds
its own inbox, triaged daily
the brains
Every model we've run it on.
repryntt is bring-your-own-key โ so we've run the same workforce on a dozen models across six providers. The engine doesn't care whose model is underneath. That's the point.
next
Benchmarks for autonomous work.
Standard AI benchmarks measure how well a model answers questions. Nobody publishes how well models complete real work end to end โ tasks with a verifiable outcome: the file written, the test passed, the email sent, the build green.
As of July 2026 our engine records the outcome of every action โ success, latency, and token cost, per model, per task type. As that data accumulates across models, this page will grow the numbers we most want to see ourselves: which models finish the highest share of real tasks, at what speed, at what cost. Same workforce, same tasks, different brains โ measured, not marketed.
coming
Task completion rate
Share of real tasks each model completes with a verified outcome โ the number that actually predicts an autonomous workforce's usefulness.
coming
Speed per task
Median latency from "task picked" to "outcome verified," per model โ how long your workforce makes you wait.
coming
Cost per completed task
Tokens spent per verified outcome โ because BYOK means this number lands on your bill, not ours.