Blog Strategy

The six months after the demo.

A prototype that impresses in a meeting and a system that survives production are different projects. The gap between them is predictable, and so is the work that closes it.

Fig. 0The demo decays; the engineered system holds

The demo went well. The model answered the questions, the stakeholders were impressed, and someone asked how soon it could go live. Six months later it still isn’t in production. Or it is, and the team spends its days firefighting.

We’ve seen this pattern often enough that it’s the reason Modulus Labs exists. The good news is that what goes wrong is predictable. The work that turns a prototype into a production system is different from the work of the prototype itself, and it can be planned.

The demo is the easy part

A prototype proves that a model can do the task on the examples you tried. That’s real progress, and it’s also the smallest part of the job. A production system has to do the task on inputs nobody tried, at real volume, within a cost and latency budget, without leaking data, while the model, the data and the business all change underneath it.

The prototype answers “can it work?” Production answers “will it keep working, and will we know when it doesn’t?”

None of this means skipping the prototype. It means running it deliberately. Time-box it to one question, such as “can a model pull these twelve fields from our invoices at 90% accuracy?” rather than “build the document system”. And before the demo, write down everything the prototype doesn’t do that production will need. Once stakeholders have seen that list, “productionize it” stops sounding like a week of polish.

Five things that break after launch

1. The long tail of inputs

Real users write differently from the team that built the demo. They’re vague, they’re in a hurry, they mix languages, they paste screenshots, and they ask about things the system was never meant to handle. Without an eval set built from real traffic, the team discovers these one complaint at a time.

2. Drift

Models get updated, documents change, and the questions people ask shift with the seasons and the business. A system that scored well at launch slowly gets worse, and no alarm goes off, because nothing is measuring quality in production.

3. Cost and latency at volume

Ten demo questions cost nothing. Ten thousand requests a day with long prompts, large retrieved contexts and a premium model can cost more than the process they replaced, and take long enough that users give up. Caching, routing simple requests to smaller models and trimming context usually fix this, but only if someone is watching.

4. Integrations and permissions

The demo read from a folder of exported documents. Production has to read from live systems, respect who is allowed to see what, write results back, and survive the day the CRM’s API changes. This is ordinary software engineering, and it’s usually most of the work.

5. Nobody owns quality

Prototypes have champions. Production systems need owners: someone who reviews failures, approves changes and decides when the system is good enough. Without one, quality becomes everyone’s concern and nobody’s job.

Is it ready? A checklist

Before we call an AI system production-ready, we want a clear yes to each of these questions, backed by evidence rather than intent:

AreaQuestionEvidence
QualityDo we know how good it is, slice by slice?An eval set built from real inputs, with current scores
ChangeCan we change it without fear?A release gate that runs the evals on every change
FailureWhat happens when the model fails?A tested fallback chain and a handoff path
SecurityWhat could a manipulated model do?Scoped tools, output checks and attack cases in the evals
OperationsWill we know when quality drops?Monitoring of quality, latency, cost and drift in production
OwnershipWho decides what “good” means?A named owner and a regular review of failures

Budget for operating, not just building

Most AI budgets cover building the system and stop at launch. But the first six months after launch are when the system meets reality, and they need engineering time: reviewing failures, growing the eval set, tuning retrieval and cost, and testing model updates. Planning for that time up front is the difference between a system that improves after launch and one that slowly decays.

If you’re choosing a partner to build an AI system, ask what happens after launch. Ask how they’ll know if quality drops, who reviews failures, and how a model upgrade gets tested before it reaches users. The answers will tell you whether you’re buying a demo or a system.

Further reading

Evals before features covers the measurement that makes the rest possible, and Prompt injection is an input-validation problem covers the security row of the checklist.