PUBLIC BETAFor launch, we turned yellow. Two products, one job: make every frontier token count.
Many point solutions, one system.
Every trick for cutting agent cost ships as its own tool: a router in one repo, a compressor in another. Prefex and ReadyBase distil 16 and 8 open-source projects into one pipeline and instrument every request that passes through it. Measured as one system, the biggest saving isn't any single trick. It's keeping the prompt cache alive.
Life of a tokenscroll or pinch to zoom · drag to pan · hover for detail
One request, drawn from the actual dispatch logic. The rule at every stop: fail open. If a stage breaks, the request goes through untouched.
01 / Where the input money goes
Caching is the product. Compression is about 2%.
Cache reads96.5%
Cache writes3.4%
Fresh input0.04%
18%of a request is tool output
×
28%of it we can claim
×
41%cut where we claim
≈
2.1%of the input stream
One changed byte near the front of a prompt re-bills everything after it, so most of the engineering is about not breaking the cache. Counted together with compression, the saving on our own traffic was about 13.8%.
02 / Built, measured, switched off
Ideas that looked good and hurt the economics.
Each came from a real project and looked like a win on its own terms. Measured on the whole system, none of them paid for itself, and the worst ones broke the cache.
Generative compaction
Saved 0.1% in 31.6 seconds. The plain extractive pass saved 42.5% in 8 ms.
Stripping old thinking
Cache reads fell from 97.6% to 38.2%. About $53 in 25 minutes.
WebSocket proxy
Upgrade attempts from real clients: zero.
Aggressive log trimming
Cut 25.9% and lost answers. Now saves 8.5% and keeps them.
03 / Bugs that taught us something
Found live, fixed, pinned by a test.
The encoder with opinions. Go sorts JSON keys and escapes <. We did it on some turns only, and a 200k-token prefix re-billed back and forth.
"All tests passed." Our summarizer said this about a failing run. Success now needs positive evidence.
Helpful history. We prepended history to a client that already sends its own. Cached tokens went to zero.
Musical chairs. Tool definitions reshuffled 69 times in 20 minutes, re-billing each time. We sort them now.
04 / What it costs you in time
Every millisecond is ours.
9.5 msmedian, 49 KB request
24.1 msmedian, 197 KB request
63.9 msmedian, 780 KB request
Against a stub that answers instantly. About 65% of it is parsing the same JSON more than once, which is the next thing to fix.
05 / Where the instrumentation leads
A local record, then a hybrid.
The system of record
Every request is logged on your own machine: model, cache reads and writes, routing decisions, what compaction did, and what it cost. ReadyBase adds what each turn changed in the code. Nothing leaves the box to build it.
Frontier to local
That record shows which kinds of turns a smaller model already gets right. It is the evidence for moving those turns to a local model, one kind at a time, while the frontier keeps the hard ones.
Offload is built: a turn you tag for a local model runs on Ollama and never quietly falls back to a paid provider. Routing turns there automatically, gated on that record, is where we are heading. It does not ship today.
That's the hood. Close it and drive.
Distilled from 16 and 8 open-source projects. The GitHub stars are theirs, not ours.