The blog
Notes on AI infrastructure
Practical writing on routing, model price-performance, and running AI in production without lighting the budget on fire.
Latest · August 18, 2026 · 6 min read
What is an LLM gateway? A production engineer's guide
A working engineer's walkthrough of the LLM gateway pattern: the request lifecycle, gateway vs direct provider calls, build vs buy, and what to check before you pick one.
Read the article
August 11, 2026 · 8 min read
Chinese open-weight models in 2026: a production guide
Open-weight models from Chinese labs now match Western flagships on many workloads at 5-20x lower prices. How the five major families compare in production — and where they still fall short.
Read the article
August 4, 2026 · 6 min read
Cut your LLM API bill by 80% without rewriting your app
Six changes, none of them a rewrite, that take an illustrative 500M-token-a-month workload from $1,750 to roughly $350 a month at current list prices.
Read the article
July 28, 2026 · 6 min read
OpenRouter alternatives in 2026: an honest comparison
OpenRouter invented this category and still has the biggest catalog. But in 2026 it's not the only sane option — here's where each gateway wins, and where it doesn't.
Read the article
July 21, 2026 · 7 min read
GPT-5.1 vs Claude vs Gemini vs DeepSeek: price-performance
Four frontier models land within a few points of each other on benchmarks — with a 60x spread in output price. What matters now is cost per solved task on your own traffic, and here's how to measure it.
Read the article