
Prefex
Same lemon, a much better squeeze.
You're paying for powerful AI. Prefex gets more useful work out of it by trimming repeat work, making better use of prompt caching and matching the job to the model.
You shouldn't need one tool for routing, another for compression, a third for small fast models and a fourth for local LLMs. Prefex brings them all into your Claude Code or Codex terminal, and you keep working the way you do.
- Stretch Claude and Codex plan limits
- Cut waste on your API bills
- Mix providers, with local models in the loop
Distils 16 proven projects with 447,644 GitHub stars between them, re-authored as a single Go binary with sidecars.
What's inside the squeeze?
Cache assist and routing do most of the saving, with compaction and session memory beside them. Savings are measured per request from provider billing fields, beyond the native caching baseline.