Kiwop Labs · quarterly series

AI agents in production

What the agents of Nexo, the platform Kiwop runs on, actually do: runs, failures, proposals signed or discarded by a person, PRs delivered by the autonomous worker, classified email and cost. A 90-day window, from 3 July 2026 to 1 October 2026. No names, no texts, no promises.

1,293agent runs
93.7 %finish without error
53.6 %client-facing proposals signed by a person
18PRs delivered by the worker
762€ of API billed

What the quarter says

  1. 28 different agents ran 1,293 times; 93.7% finished without error. The median run takes 88 s and the slowest 10% exceed 175 s.
  2. Agents proposed 500 task comments; 168 were client-facing and a person signed 90 (53.6%). The rest were discarded or are still waiting. Signing takes a median of 78.2 hours: human review is not a formality.
  3. The autonomous worker claimed 25 delegated tasks and delivered 18 PRs, with a median of 0.7 hours between claiming and delivering. The director reviewed 155 tasks and delegated only 23.
  4. In email, the classifier handled 216 inbound threads and created 31 tasks; the guardian reviewed 1,247 outbound messages and blocked 111. People accepted 68.8% of its suggestions.
  5. Cost: 762 € billed via API in the window, plus 1,736 € of subscription usage valued at API prices. The CLI on the server consumed 3,284 $ at list price, 32% of it by the autonomous worker.

Proactive agents

The lookouts: agents that fire on their own (by cron, by event or by another run), read the project state and propose something or raise an alert. Each run stores status, duration and alert level.

MetricValue
Runs1,293
Completed1,212
Failed69
Success (%)93.7
Active agents28
Median duration (s)88
p90 duration (s)175
By alert levelwarning: 230 · ok: 910 · sin nivel: 81 · critical: 72

Proposals signed by a person

A comment written by an agent is born pending and invisible until the person whose name it carries signs it in the web app. There is no way to sign it via API. This is the real human acceptance rate, not an estimate.

MetricValue
Proposed comments500
Internal notes332
Client-facing168
Signed354
Signed (%)70.8
Client-facing, signed90
Client-facing, signed (%)53.6
Hours to signature (median)78.2

Autonomous worker and task director

The director reviews every new task and decides whether to delegate it to the worker (a Claude Code on the server) or leave it to a person. The worker claims, works in an isolated worktree and delivers a PR. Nobody merges for it.

MetricValue
Tasks claimed by the worker25
PRs delivered18
Hours from claim to delivery (median)0.7
Worker passes48,987
Average queue0
Tasks reviewed by the director155
Delegated to the worker23
Handed to a person10

Email: classifier and guardian

The classifier reads inbound email and decides whether it is a task, a reply or noise, with its confidence. The guardian reviews outbound email before it leaves and returns ok, suggestion or block. Both are corrected by a person.

MetricValue
Threads classified216
By typereply: 130 · ignore: 45 · task: 41
By confidencesin dato: 157 · high: 40 · low: 19
Classifier failed (%)4.6
Median latency (ms)10,249
With a decision124
By decisionthread_duplicate: 19 · sin decidir: 73 · mirrored: 28 · auto_attached: 65 · auto_created: 31
Tasks created31
Reviews1,247
By verdictsuggest: 93 · block: 111 · ok: 1,043
Suggestions accepted (%)68.8
Median latency (ms)7,849
Suggestions per review (mean)0.24

Queries to the house criterion

Agents and the team ask the “brain” (the documented criterion of the company) before deciding. If the answer does not reach the minimum similarity, it is escalated to a person.

MetricValue
Queries154
Escalated to a person66
Escalated (%)42.9
Rated0
Useful (%)–
Median similarity0.54

Lead agent

Automatic follow-up of inbound leads: drafts a person reviews, automatic stop if they are not reviewed, and discarding of vendors and noise.

MetricValue
Sequences75
By statuswaiting: 4 · stopped: 63 · replied: 6 · exhausted: 2
Follow-ups sent11
By stop reasonactiva: 12 · el lead ya no existe: 2 · not_business: 9 · parada manual por Josep Purroy: el correo del lead rebota (554 5.7.1): 1 · reconciliado: la conversación ya existía en el buzón: 1 · borrador sin revisar durante 4 días: 4 · borrador sin revisar durante 7 días: 1 · proveedor haciendo outreach: fuera del pipeline comercial: 4 · clasificado como «Otro»: no se contacta: 18 · borrador sin revisar durante 6 días: 1 · ya le contestaron a mano desde josep@kiwop.com: 1 · clasificado como «Proveedor»: no se contacta: 19 · particular que quiere formarse: con el correo del curso basta: 1 · clasificado como «Prácticas»: no se contacta: 1

AI cost

Per-call metering of everything that goes through Nexo’s AI gateway. We separate what is billed via API from what goes through subscriptions (valued at API prices so it can be compared, but not paid per token).

MetricValue
Calls11,500
Input tokens564,418,480
Output tokens10,685,590
Billed via API (€)761.81
Via subscription, API equivalent (€)1,736.09
By provideropenai: 1,008 · anthropic: 10,492
Distinct models13
Median latency (ms)12,136
Cache reads (%)0
FeatureCallsInput tokensOutput tokensCost (€)
Proactive agents6,307324,460,6288,697,3971,524.84
Chat2,27598,426,113489,634419.55
Email guardian1,60578,452,290512,331240.29
Email classifier78831,652,077432,698172.74
Project chat17115,614,418220,95555.18
Outbound email24610,067,22488,49752.23
Academy774,324,947117,37124.98
Newsletter15758,60282,1935.85
SEO content13561,32542,0781.64
CRO2100,8562,4360.57
operations_slack_audio1000.03

And the other ledger: CLI sessions on the server (the autonomous worker, the relay and the platform binary), measured from transcripts and valued at list price.

OriginSessionsTurnsList cost ($)
Platform binary5,36210,6181,168.64
Autonomous worker5496,1541,059
Claude relay2,6203,5611,032.92
Brain relay1,1952,37122.89
Unclassified7100.16
Total9,73322,7143,283.61

The crons that hold it up

Everything above runs on scheduled tasks. This is their full history in the window: how many, how many were skipped on purpose and how many failed.

MetricValue
Runs259,134
Distinct scheduled tasks256
By outcomeskipped: 17,519 · failed: 47 · running: 1 · ok: 241,567
Failed (%)0.02
Median duration (ms)1,872

Why we publish this

  • Because almost everything written about agents in production is a promise. These are counters from a platform that has been running a real agency for months, with its failures and its discards.
  • Because the human signature rate is the number that matters most and is published least: how many of an agent’s proposals survive the person who reads them.
  • Because the real cost (two ledgers: API and subscription) is what decides whether an agent pays off, and almost nobody shows it.

Method

  • One read-only SQL query on Nexo’s database, published in full. It runs every quarter over a 90-day window and produces aggregates only.
  • Nothing identifiable: no clients, no projects, no texts, no ids. Distributions and medians.
  • The data comes from one company (Kiwop) and one platform (Nexo). It is not a market sample; it is one complete real case.
  • The definitions (what a run is, what a signature is, what a delivered PR is) live in the query and in the Labs documentation.

Limits

  • The window spans changes to the platform itself: new agents, cache pricing fixes, delegation rules. A jump between quarters can be a change of ours, not of the world.
  • Subscription cost is valued at API prices for comparison; it is not money paid.
  • Unsigned comments include those discarded and those still waiting: we do not tell them apart.
  • A run’s duration includes waiting for the API.

Open data

How to cite

Kiwop Labs (2026-10). AI agents in production: aggregated telemetry from Nexo. https://www.kiwop.com/en/labs/ai-agents-in-production

Data under CC BY 4.0. Query and method open.

← Back to Kiwop Labs

Let's talk.

Initial technical consultation

AI, security and performance. Diagnosis with phased proposal.

  • NDA available
  • Response <24h
  • Phased proposal

Your first meeting is with a Solutions Architect, not a salesperson.

Talk to an architect