dimaforcepush

RESEARCH

Anatomy of a quota burn

How one agent session spent the last 110 million tokens of a weekly quota overnight — and why no per-turn budget noticed.

METHOD
Read-only SQLite snapshots and agent logs. Token counts and timings only, no message content.
WINDOW
21 Sep 06:00 → 25 Sep 07:05 UTC, 4.05 days.
RESULT
A supervisor that answers routine worker prompts without calling a model.
weekly quota used, 21 → 25 Sep
22% → 100%
tokens in the window, 93% of them cache reads
406M
background wake-ups of one session in six hours
211
tokens of context re-read on every wake-up
~240K
01

Most of it was an ordinary week

The first suspicion was one bad night. The logs said otherwise: about three quarters of the quota was gone before that night, spent on ordinary daytime work across three agent profiles.

Tokens per day, all profiles
LabelValue
21 Sep96M
22 Sep50M
23 Sep51M
24 Sep128M
25 Sep, 00:00–03:0669M · the night

The night added the last ≈110M — enough to finish the quota, not enough to explain it alone.

02

One session, woken 211 times

The session that tipped it over was supervising a Claude Code worker. The pattern was: start a background watcher, end the turn, get woken when the worker changes state. The worker was in plan mode and asked for approval every one or two minutes. Every question ended the watcher, and every end of the watcher woke the supervisor with a fresh turn.

API calls per hour, 24 Sep 19:00 → 25 Sep 01:05 UTC
LabelValue
19:0044
20:00199
21:00154
22:00118
23:00158
00:00116
01:0019 · quota gone

Each wake-up ran three or four calls — far below the per-turn limit of 90. No single turn was expensive, so no per-turn budget ever fired. The loop lived across turns.

03

The context never shrank

Every one of those calls re-read the whole conversation: the per-call context grew from about 100K to about 243K tokens and stayed there. Compression ran 37 times and made no progress 37 times; each failed pass also duplicated the transcript on disk.

04

What changed

The root cause was a skill change two weeks earlier: workers moved from an auto-approve permission mode to plan mode, which prompts for each shell command — and each prompt woke the supervisor.

The fix is a small supervisor script that sits between the worker and the agent. It answers routine prompts itself, keeps a deny-list for pushes, deploys and destructive operations, and wakes the agent only for a ready plan, a real question, a finished job, a denial, a stuck worker or a budget cap. It has unit tests, and the old watcher is kept for rollback.

  1. 01Workerasks for approval, or changes state
  2. 02Supervisorroutine? answer it without a model
  3. 03Agentwoken only for plans, questions, results, denials