A pattern that recurs: a team builds a chat-shaped feature on top of their product. The model layer is excellent. The retrieval layer is plausible. The system gives confident answers that are subtly, persistently wrong, and the team can’t tell why because they can’t trace what the model was actually given.
The model is not the problem. The model is doing exactly what it was asked to do. The problem is everything underneath it: the data that fed retrieval was duplicated across four upstream systems, none of which agreed; the lineage went cold three hops back; the retrieval index was rebuilt monthly from a snapshot that was already stale by the time the index was warm. The model was answering with confidence about state that didn’t exist.
This post is the field guide we wish more teams had read before they started. It is not about LLMs. It is about the plumbing AI products forget to plan for.
Ingest is where the lies start
Every AI product has an ingest layer. Most of them have not been audited as ingest layers. They were built as “the way we got data into the index for the demo”, and they’ve been tolerated since because the demo worked.
The audit is short and uncomfortable. For each source feeding the index:
- Where does the data physically come from? Name the system, the API, the dump, the export job.
- Who owns it? Not the system — the data. The team whose name is on the change ticket when the schema moves.
- How fresh is it? Sub-second, hourly, daily, weekly, “depends on whether the cron is healthy”.
- What’s the failure mode? When this source goes silent, does the index degrade gracefully or does it return stale answers as if they were live?
The output of that audit is usually humbling. Two of five sources are flowing through an export job nobody owns. One source is a quarterly snapshot the team forgot was quarterly. The “real-time” source has a P95 lag of eleven minutes that nobody had measured. The model is answering about state that is, on average, days off.
The fix isn’t “fix the sources”. The fix is to acknowledge what they are, design the product around their actual freshness profiles, and surface the limitation in the answer. “The most recent data point in this answer is from yesterday’s batch.” That sentence, in the UI, is the difference between a feature users trust and a feature users learn to second-guess.
Lineage is the only thing that survives an incident
When an AI feature returns the wrong answer, the question every business asks is the same: where did that come from? If the answer is “the model said it”, you don’t have a product. You have an oracle. Oracles are not auditable.
Lineage is the discipline of answering the question deterministically. For every retrieved chunk in every answer: which source did it come from, when was it ingested, what version of the embedding model produced its vector, what version of the chunking strategy split it. Stored alongside the response, queryable when a user complains.
# Sketch of the metadata we attach to every retrieved chunk.
# Stored with the response. Queryable from the support tool.
@dataclass
class RetrievedChunk:
text: str
score: float
source_system: str # "salesforce", "confluence", "pim"
source_id: str # natural key in the source
source_revision: str # commit, version, ETag, etc.
ingested_at: datetime
chunker_version: str
embedder_version: str
index_version: str
document_owner: str # team that owns the upstream record
That dataclass looks unglamorous. It is the difference between a support engineer being able to say “the model retrieved this chunk, from this source, ingested at this time, from this revision” and the support engineer having to say “we don’t know”. The first sentence closes tickets. The second is how products lose their users.
Lineage is also the precondition for ever swapping the model out. Teams that don’t track which embedder version produced which vector can’t reindex incrementally; they have to nuke the index and rebuild on every change, which means in practice they never swap embedders, which means they’re locked into whatever was current the day they shipped.
Governance, before the regulators ask
The “AI feature” route to production usually skirts data governance. The reasoning is that retrieval-augmented generation isn’t really a new use of data — it’s just a clever search. The reasoning is wrong. Retrieval makes implicit decisions about access: a user is now seeing a synthesized answer that may include text from documents they have direct read access to and text from documents they don’t.
The governance question that matters: at retrieval time, does the index respect the same access controls as the upstream source? In most early implementations we audit, the answer is no. The index is built once with elevated permissions; everyone querying it sees the same retrieved chunks, regardless of their entitlements. That’s a data leak waiting to happen, and once it has happened, the remediation is expensive — you can’t unsay what the model said.
The right shape, before the regulator asks: filter retrievals by the calling user’s entitlements; record which entitlements were applied; have a defensible audit trail. For workloads under HIPAA, NHS DSPT, or PCI-DSS, this is mandatory. For the rest, it is best practice that becomes mandatory the first time anything goes wrong. Skipping it before launch is exactly the kind of choice that makes the post-launch hardening conversation unwinnable.
Retrieval makes implicit decisions about access. The index respects upstream entitlements, or it leaks them.
Retrieval is not just embeddings
Vector search has earned its place. It is also not the entire retrieval story, and treating it as such produces brittle systems.
The retrieval layer that holds up in production is hybrid. Lexical search for high-precision queries with proper-noun anchors (“show me the Q3 SOC 2 report”). Vector search for semantic queries with no obvious anchor (“what are our incident commander responsibilities”). Filters for entitlement, recency, and source quality. Re-ranking — by a smaller, faster model — to combine signals before the top-k goes to the LLM.
Each of those layers has parameters that need to be tuned against real query traffic, not against a developer’s intuition. The retrieval evaluation rig — a set of representative queries with expected sources, run on every change to the retrieval pipeline — is the most underrated piece of infrastructure in AI plumbing. Teams that build it ship faster and regress less. Teams that don’t ship a feature whose quality depends on whoever last tweaked the chunker’s mood that afternoon.
The chunking strategy is the product
If the chunking strategy is “split documents into 1000-token windows with 200-token overlap”, the team has not chosen a chunking strategy. The team has accepted the default in the first tutorial they read.
Chunking is the most impactful single decision in the retrieval pipeline. It determines what the model can possibly see. A document chunked badly — across section boundaries, in the middle of a table, separating a heading from its content — will retrieve poorly even with a perfect embedder.
The chunking strategies that work in production are document-type-aware. Markdown chunked at headings. PDFs chunked with structural awareness, not at character offsets. CSVs chunked row-coherently with the header repeated. Code chunked at function boundaries. Each requires a parser. Each parser is unglamorous engineering work that is invisible from the demo and load-bearing in production.
The team that treats chunking as a product decision — versioned, evaluated, owned — ships better answers. The team that treats it as a config knob ships answers that are sometimes good and sometimes mysteriously poor.
The model layer is swappable. The data layer is what you keep.
The hardest lesson of the last two years, for teams that took it: the model is the part of the stack that will change the most often. Frontier models will get better. Open-weights alternatives will close the gap on specific tasks. Cost curves will move. Inference providers will bid against each other. A team that has bet its architecture on a specific model is going to be doing migration work, repeatedly, on a cadence the model providers control.
The data layer is the part of the stack the team owns. Ingest pipelines, lineage records, governance controls, retrieval evaluation rigs, chunking strategies, the actual indexed corpus. None of that changes when the model does. All of it determines whether the next model is better in your hands or just on the leaderboard.
The strategic implication is what informs how we sequence the work: build the data layer to be model-agnostic from the start. The retrieval pipeline returns chunks. The orchestration layer formats those chunks into a prompt. The model is one swap away. Embedders are versioned and reindexable. The vector store is one of three in the team’s reach. The whole pipeline can be re-pointed at a different provider in a sprint, not a quarter.
What “production-ready” actually means here
A working AI feature is not the same as a production AI feature. The list that distinguishes them is short and depressingly often skipped:
- Lineage on every retrieved chunk, queryable from a support tool.
- Entitlement filtering at retrieval time, with audit logs.
- A retrieval evaluation rig that runs on every pipeline change.
- A defined freshness SLO per source, surfaced in the UI when stale.
- A reindexing path that doesn’t require a full rebuild.
- A way to roll back a chunker, embedder, or model version without an outage.
- An incident playbook for “the model is confidently wrong about a known fact”.
Each of those is the kind of work that doesn’t show up in the demo. Each of them is what the team will be glad they did the first time a customer escalates an incorrect answer.
The honest sequence
The order this work has to happen, regardless of how the deck reads: data layer first, retrieval second, orchestration third, model fourth. The teams that do it in that order ship slower in week one and faster in month four. The teams that do it in the other order ship a demo in week one and have to redo the foundation in month four, with users now watching.
There’s a more general version of this argument we keep making: AI features fail at the layer that’s least glamorous to fix. If your retrieval is brittle and your lineage is implicit, the cleverness of the prompt isn’t going to save you. It’s the plumbing. It always was.
If you have an AI feature in production whose answers are subtly drifting, the diagnosis is almost certainly downstairs from the model.