Many point solutions,
one system.

Every trick for cutting agent cost ships as its own tool: a router in one repo, a compressor in another. Prefex and ReadyBase distil 16 and 8 open-source projects into one pipeline and instrument every request that passes through it. Measured as one system, the biggest saving isn't any single trick. It's keeping the prompt cache alive.

Life of a token scroll or pinch to zoom · drag to pan · hover for detail
Phase 1, the client: whatever tool made the API call (Claude Code, Codex, Cursor, aider). It sees one endpoint and none of what follows. ① CLIENT Phase 2, the proxy: everything in this band runs on your machine, inside one binary. Nothing crosses the network until the gateway on the right. ② PREFEX PROXY · runs locally, inside one binary MAIN PIPELINE ── fail-open at every stage: internal errors log and forward unchanged Sharpen: an optional prompt-refinement pass you run yourself, before a request exists. Not on the live path. ✎ SHARPEN (CLI) prompt refine, before sending prefex wrap: launches one tool through the shared daemon with its base URL set for that process only. Your global settings stay untouched. ⇢ prefex wrap (CLI) one tool, one process Vault: an opt-in, separate proxy for CLI and tool HTTP calls. It swaps sealed references for real secrets on the way out and scrubs them from anything logged. Off by default, and not on the LLM path. 🔒 VAULT + MITM opt-in · own port · off by default seals tool secrets, not the LLM path Session loop: the API is stateless, so the whole conversation is re-sent every turn. The unchanged front bills at a tenth when the cache holds, which is why every byte there is guarded. session context re-sent every turn (cached: 0.1× rate) Coach loop: findings point at settings changes, and once applied they shape every later session. The slowest, widest loop here. coach finding, then a settings change, then every later session Return trip: the reply streams back through the same path. PII placeholders are restored, an optional stream guard can cut a blocked answer at a block boundary, and logging happens off the critical path. ← response streams back · PII restored · logged off the critical path Recall: a lossy compaction keeps its original and leaves a marker. The model can ask for the original back, and we count how often it does. recall Request: an Anthropic- or OpenAI-shaped call arrives on localhost. REQUEST client call Auth: works out who you are and which team pays. AUTH key → dev/org Channel: tags the source tool from its user agent and path. CHANNEL tool source Session: loads prior turns. Both loops below come back in here. SESSION prior turns Cache: keeps the front of the prompt byte-identical. Tool order is canonicalized and the body is re-serialized the same way every turn, because a single flipped byte re-bills a 200k-token prefix at the write rate. CACHE prefix? tool order + bytes pinned ● HIT · 0.1× read MISS · 1.25× write Router: signals, score, tier, capability. Stays inside one model family; pins and named overrides switch it off. ROUTER model? ▽ expanded below Shaper: an optional verbosity instruction, frozen per session so a settings edit never moves an open session's prefix. Measured about 11% fewer output tokens, with a holdout arm to prove it. SHAPER sys-prompt inject Compress: twelve content types, lossless where it can be, lossy only with a recoverable original. Anything carrying a protected marker is declined whole. COMPRESS ▽ expanded below PII and guardrails: guardrails judge what you sent before we touch it; PII is swapped for placeholders after compaction and restored in the reply, tool arguments included. PII · GUARD judge · mask Forward: the last thing done to the request before it leaves. FORWARD stream upstream ③ MODEL GATEWAY · leaves the machine Model gateway: the only hop we do not control. Anthropic by default, plus OpenAI, Gemini, Bedrock, Vertex, OpenRouter or a local Ollama. 🌐 MODEL GATEWAY Anthropic · OpenAI · Bedrock · Vertex · Ollama the one hop we do not control Log: one row per request in local SQLite, written after the response has started streaming. LOG SQLite + async verbosity, frozen per session holdout arm · ~11% output cut guard judges what you sent PII masked, restored on return ROUTER · signal → score → tier → capability SIGNAL EXTRACTION Prompt length, log-scaled.prompt len Input tokens, log-scaled.input tokens Two or more code fences.code fences Question marks, capped.questions Words that usually mean hard work (architect, security, debug) push up; words that usually mean chores (rename, format, quick) push down.strong/weak terms Weighted sum of the signals, clamped to 0 to 1. SCORE 0 to 1 Optional: a small local classifier blends a semantic score in. If it is unreachable, the heuristic runs alone. ⚡ ONNX encoder (semantic blend, optional) MODEL LADDER (5 tiers · one family · pins win) Strongest tier. Default: Opus 5.5.Strongest Strong tier. Default: Sonnet 5.Strong Medium tier. Default: Sonnet 5.Medium Lite tier. Default: Haiku 4.5.Lite Cheapest tier in your ladder.Cheapest ≥.75 ≥.55 ≥.35 ≥.15 <.15 CAPABILITY FLAGS (per chosen model) Adaptive thinking, where the model supports it.adaptive thinking Some models refuse a request that names no thinking config.thinking required Effort, where the model accepts it.effort Some models gain little from long reasoning, so we do not ask for it.low-CoT model OFFLOAD BRANCHES & GATES Named local models run on Ollama. They never quietly fall back to a paid provider.Ollama offload (local) Gemini SDK calls take their own path.Gemini inbound (SDK) Above 80k input tokens the router stops downgrading, to avoid a smaller context window.skip downgrade > 80K tok COMPRESS · content router → tier → gate → CCR CONTENT-TYPE ROUTER (12 types) JSON arrays become lossless tables, checked by a round trip.JSON Logs: head, tail, errors with context. A middle line is dropped only if its template already appeared.Log grep and ripgrep output grouped by file, losslessly.Search Source code: bodies elided by a real parser, signatures kept verbatim, line numbers intact. A file that would save under 20% is left alone.Code Free text: a native extractive pass, then an optional extractive model.Text Tool results with no dedicated filter fall through to the text path.Diff·Config·CSV SAFE · lossless generic Read Grep Glob Fetched pages: markup dropped, every visible character kept.HTML · Path BALANCED · lossy Minified one-line blobs keep their head and tail.Read.balanced Long search results collapse repeated directories.Grep.balanced Bundled summaries for git, cargo, pytest, npm, docker, terraform, go test and eslint. A summary may only say 'passed' with positive evidence.git · cargo · pytest · npm · docker · terraform · go · eslint ReadyBase gate: a repeat read is skeletonized only when the file is known, low blast radius and tested. Anything else stays whole. ReadyBase gate: known ∧ ¬high-blast ∧ tested → skeletonize · else → bare reference Protected markers: a block carrying a recall or ReadyBase marker is declined whole, at the entrance and again at the commit point. protected markers: decline Lossy results keep the original and carry a marker, so they can be recalled. lossy → <<ccr:HASH>> marker An optional extractive model on a local sidecar. Off by default; on any error the text passes through. ⚡ extractive model (optional) FEEDBACK & BACKGROUND ── async, never blocks the response Recall store: the originals behind every lossy compaction. CCR STORE recoverable originals Status line: per-directory turn stats for your terminal. STATUS LINE per-directory Prewarm: before a restart, replays each session with max_tokens 0 to pre-pay the cache write while nothing is broken. PREWARM pre-pay writes before restart · max_tokens=0 Advisor: an optional local model that can nudge routing and add sticky guidance. Only acts above a confidence floor. ADVISOR optional local model acts only above a confidence floor Recorder: every request's tokens, cost, route and compaction, in local SQLite. Drives the 'Saved by prefex' figure. RECORDER local SQLite cost/savings · routing/compress · session/project drives the Saved by prefex figure Setup health scan, cached for 5 minutes. setup health · 5m Learns per-repo verbosity signals, every 30 minutes. output learn · 30m Coach: evaluates findings against fresh telemetry. COACH inline · 5s Billing watch, team tier only. billing watch · team · 24h Fail open: on an internal error, log it and forward the request unchanged. // every stage above fails open: on an internal error, log it and forward unchanged Session A: Claude Code Session B: Codex Session C: Cursor Session D: aider
One request, drawn from the actual dispatch logic. The rule at every stop: fail open. If a stage breaks, the request goes through untouched.

01 / Where the input money goes

Caching is the product.
Compression is about 2%.

Cache reads96.5%
Cache writes3.4%
Fresh input0.04%
18%of a request is tool output
×
28%of it we can claim
×
41%cut where we claim
≈
2.1%of the input stream

One changed byte near the front of a prompt re-bills everything after it, so most of the engineering is about not breaking the cache. Counted together with compression, the saving on our own traffic was about 13.8%.

02 / Built, measured, switched off

Ideas that looked good
and hurt the economics.

Each came from a real project and looked like a win on its own terms. Measured on the whole system, none of them paid for itself, and the worst ones broke the cache.

Generative compaction

Saved 0.1% in 31.6 seconds. The plain extractive pass saved 42.5% in 8 ms.

Stripping old thinking

Cache reads fell from 97.6% to 38.2%. About $53 in 25 minutes.

WebSocket proxy

Upgrade attempts from real clients: zero.

Aggressive log trimming

Cut 25.9% and lost answers. Now saves 8.5% and keeps them.

03 / Bugs that taught us something

Found live, fixed, pinned by a test.

  • The encoder with opinions. Go sorts JSON keys and escapes <. We did it on some turns only, and a 200k-token prefix re-billed back and forth.
  • "All tests passed." Our summarizer said this about a failing run. Success now needs positive evidence.
  • Helpful history. We prepended history to a client that already sends its own. Cached tokens went to zero.
  • Musical chairs. Tool definitions reshuffled 69 times in 20 minutes, re-billing each time. We sort them now.

04 / What it costs you in time

Every millisecond is ours.

9.5 msmedian, 49 KB request
24.1 msmedian, 197 KB request
63.9 msmedian, 780 KB request

Against a stub that answers instantly. About 65% of it is parsing the same JSON more than once, which is the next thing to fix.

05 / Where the instrumentation leads

A local record,
then a hybrid.

The system of record

Every request is logged on your own machine: model, cache reads and writes, routing decisions, what compaction did, and what it cost. ReadyBase adds what each turn changed in the code. Nothing leaves the box to build it.

Frontier to local

That record shows which kinds of turns a smaller model already gets right. It is the evidence for moving those turns to a local model, one kind at a time, while the frontier keeps the hard ones.

Offload is built: a turn you tag for a local model runs on Ollama and never quietly falls back to a paid provider. Routing turns there automatically, gated on that record, is where we are heading. It does not ship today.

That's the hood.
Close it and drive.

Distilled from 16 and 8 open-source projects. The GitHub stars are theirs, not ours.