Blog
Notes on adaptive inference.
What we’re learning about evaluation, inference, and learning from experiments. Every figure carries its source, and every post ends with what would change our mind.
-
Two loops
The part that learns is slow and expensive. The part that serves has to be fast and predictable. The interesting engineering is the boundary between them.
-
Cost per successful task, not cost per token
Cheap failures are the most expensive thing you can buy.
-
You can’t prove savings without a counterfactual
A falling bill doesn’t show that the thing you bought worked. Prices fell too, and the work changed.
Review build — these posts are drafts and haven’t been published.
Work with us
Bring a task worth improving.
We’re looking for design partners with repeated AI workloads, measurable outcomes, and a clear reason to improve performance. Work with us to define the problem and test what better execution could look like.
Working on evaluation, inference, or learning from experiments? We’d like to hear from you.
Discuss research collaborationReview build — contact address not yet connected.