Your demo worked. That was the easy part.
It ran on ten hand-picked examples, the ones you knew would land. Production is the long tail of real inputs: the typos, the edge cases, the questions nobody on your team thought to test. The demo proves the model can do the thing once. Production asks whether your system does the thing a million times without falling over, leaking money, or answering from last month’s data.
Four things break when you ship AI. None of them are model problems.
No evals, no improvement
You cannot improve what you do not measure. Without evals, every change to the prompt, the retriever, or the model is a guess. It felt better in the demo. The users say it feels worse. Nobody can prove either side, so the loudest voice in the room wins, and the loudest voice is usually wrong.
Evals are the unit tests of an AI system. A fixed set of real questions with known-good answers, run on every change, with a number that goes up or down. This is not research. This is the same discipline you apply to a dbt model: you do not merge without the tests passing.
Stale index, wrong answers
Yesterday’s docs answer today’s questions. The retriever did its job perfectly and returned the top chunk, which was written before the policy changed, before the price changed, before the feature launched. The model cited it beautifully.
This is the same failure from the RAG post, and it keeps showing up because teams treat the index as infrastructure instead of data. Infrastructure you set up once. Data you refresh on a schedule, with an SLA, with alerts. The index is data.
Latency times cost
The big model at real traffic blows two budgets at once. The dollar budget: per-token pricing multiplied by production volume is a number that surprises everyone the first month. The latency budget: your p99 goes from acceptable to embarrassing the moment real users hit the endpoint concurrently.
Demos hide both. One user, one request, no concurrency, no bill. Production multiplies everything by the traffic you actually have, and the model that felt instant in the demo is suddenly the slowest box in your system and the most expensive line on the invoice.
The inputs nobody tested
Edge cases will humble your prompt in days. The demo inputs were clean sentences from cooperative colleagues. Production inputs are pasted error logs, screenshots described in broken English, questions in three languages, and the occasional user testing whether the bot will say something it should not.
You cannot prompt your way out of the long tail. You can only measure it, sample it, and keep feeding the weird ones back into your eval set. The tail is the product. The demo was the trailer.
Shipping AI is data engineering
Pipelines, measurement, iteration. Evals are tests. Freshness is an SLA. Latency and cost are capacity planning. The long tail is data quality work. Every fix on this page is a thing data engineers already know how to do.
The model is the smallest part of the system. The demo proved the smallest part works. Now build the rest.
This post started as an Instagram post →
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.