The demo takes an afternoon. Making it reliable enough to put in front of customers is the other 95%.
Every team we work with has built an impressive LLM demo in a weekend. Far fewer have one in production that they trust. The gap between those two states is almost entirely unglamorous engineering, and it's worth knowing what's in it before you commit to a roadmap.
Start with the evaluation, not the prompt
The first thing we build on an AI feature is not the feature — it's a set of thirty to a hundred real examples with known-good answers, and a script that scores the current system against them. Without it you have no way to tell whether your prompt change on Tuesday made things better or quietly broke a category of input you weren't thinking about.
Assembling those examples is tedious and requires someone who knows the domain. It's also the single highest-leverage day of work on the whole project. Run the suite in CI, and treat a regression the same way you'd treat a failing unit test.
Retrieval is the hard part, not generation
For most business use cases — document Q&A, support answers, clause search — output quality is dominated by whether the right context made it into the prompt. When we're called in to fix a RAG system that "hallucinates", the problem is almost never the model. It's:
- Chunking that splits meaning. A clause cut in half retrieves as nonsense. Chunk on document structure, not a fixed token count.
- Pure vector search where keywords matter. Product codes, names and legal references need lexical matching. Hybrid search plus a reranker beats embeddings alone in nearly every evaluation we've run.
- No handling for "the answer isn't here". If nothing relevant is retrieved, the system must say so. That behaviour has to be designed and tested, not hoped for.
Design the failure state first
The model will be slow, wrong, or unavailable at some point in every week of production. So the interface question isn't "how do we show the answer" — it's "what does this screen do when there isn't one". Stream tokens so latency is visible rather than mysterious. Cite sources so users can verify rather than trust. Make correction cheap and obvious. And always have a path to a human for the cases that matter.
Put a number on every request
Track tokens and cost per request from day one, broken down by feature. Costs that look negligible in testing become a line item at scale, and the usual culprit is a retrieval step quietly stuffing 8,000 tokens of context in to answer a one-line question. Set hard per-user and per-org limits before launch, not after the first surprising invoice.
The maintenance question nobody asks
Models get deprecated. The version you launched on will be retired, and the replacement will behave differently — usually better overall, occasionally worse on your specific task. If you have an eval suite, migrating is an afternoon: run it against the new model, look at what moved, adjust. If you don't, it's a month of anxiety and customer reports.
That's the real argument for all of this scaffolding. It isn't rigour for its own sake. It's what makes the feature something you can change later without fear.