Anant Srivastava Says Your Enterprise AI Architecture Was Accumulated, Not Designed – finance.biggo.com

If your company’s internal AI assistant starts recommending products that don’t exist, the failure probably won’t trace back to a single catastrophic decision. It will trace back to a product manager quietly tweaking a prompt, a support engineer adding documents to retrieval, and a machine learning team fine-tuning on six-month-old tickets — each move reasonable in isolation, and nobody owning the whole. That, according to Anant Srivastava, speaking on the AI Engineer podcast, is how enterprise AI architectures actually get built: they are accumulated, not designed.
Srivastava, who presents himself as a practitioner who has made these mistakes firsthand, argues that most AI teams never consciously decide where knowledge lives. Instead, they inherit a system that grew through uncoordinated incremental moves. His central diagnostic is a single question applied to every piece of knowledge the system holds: is it small, stable, and about behavior? Is it large, current, and citable? Or has it stopped changing and converged? The answer determines whether it belongs in the prompt, in memory, or in the model’s weights.
Srivastava rejects the metaphor most teams implicitly use — that prompt engineering comes first, then retrieval-augmented generation, then fine-tuning, as escalating stages of sophistication. “This is not a ladder that you climb,” he said. “These are three different tools for three different jobs.”
The prompt’s job is tone, persona, and behavior. His example is a customer support agent that keeps a professional tone, offers human escalation after three failures, and provides concrete next steps. None of that changes per query or per user, so it belongs in the prompt. What doesn’t belong there is facts. Stuffing a product catalog into the prompt means paying token costs for context the model doesn’t need on every call, and large token loads trigger the well-documented lost-in-the-middle problem where models ignore information buried in long contexts.
Memory, by contrast, is where knowledge that is current, large, and citable lives. Srivastava uses “memory” to cover two things: agent memory about the person interacting with the system, and external enterprise memory retrieved via RAG, such as refund policy documents. The external store is managed by an application, not the agent itself.
Access control is a hard requirement for memory placement. If some users must not see other users’ data, that knowledge has to live in memory so the system can scope who sees what. His example is a code assistant with access to an organization’s entire repository estate. You wouldn’t fine-tune a model to learn the code, and you wouldn’t stuff the codebase into a prompt. Instead, you build retrieval with code-aware chunking — using AST parsing to split code at function boundaries rather than arbitrary character cuts — and you denormalize metadata such as which repo a chunk belongs to and which users are allowed to commit. Without that filtering, Srivastava warns, you create what he calls a “RAG mush”: many chunks returned from the database that confuse the model rather than ground it.

The quietest failure Srivastava describes doesn’t involve escalation at all. Teams not actively fixing a problem still make placement decisions daily. “Every edit that you make to the prompt decides what behavior lives in the prompt,” he said. “Every document that you index decides what knowledge is held in the memory, and every training sample that you identify will bake something into the weights.”
His extended example is an internal support assistant. Within a few weeks of release, a product manager edits the prompt to fix tone. The support team adds refund policy documents to retrieval. An engineer shoves the product catalog into the prompt because the agent isn’t answering correctly. Then the ML team fine-tunes on support tickets from up to six months prior.
Each decision looks fine in isolation — Srivastava concedes that one of them, the catalog in the prompt, is not — and he admits he has made similar decisions without much thought. The failure only becomes visible when a new product catalog launches. The team updates the prompt copy, and the agent still returns product names that don’t exist. The old catalog leaked into the weights during fine-tuning on stale tickets, and no one knew to look there.

Srivastava is explicit that fine-tuning has legitimate uses — but they are narrower than most teams assume. The right use case is a hard, ambiguous problem such as content moderation or claims processing. Start with a frontier model making recommendations, have humans correct it — inconsistently at first — and over time patterns emerge: a stable center plus a contested edge where humans are still needed. Human override rates flatline, and the stable center becomes fine-tuning material, with drift monitoring attached because if the center moves, the fine-tune breaks.
The wrong use, he says, is fine-tuning to fix retrieval problems. His cautionary example is an internal document assistant where the team fine-tunes on runbooks and procedures because the model isn’t giving right answers. That was a retrieval problem — the model needed the right chunks, not new weights. Fine-tuning instead produces a model that serves stale information. Runbooks and procedures are facts, and facts belong in external memory.
He draws the same line with medical claims coding. ICD-10 codes, roughly 70,000 of them, are knowledge you would not fine-tune: the codes change, and you would need a massive training set. What you fine-tune is the reflex — understanding the input format and selecting across formats.
The two acceptable reasons to fine-tune, in Srivastava’s framing, are capability and cost. Capability is rarely the issue with frontier models. Cost usually is: a small fine-tuned model can run high volume cheaply if the work has a stable pattern. That makes the retrieval-versus-weights boundary, at least partially, a budget question rather than a purely technical one.
The three layers don’t just sit side by side — they circulate into each other. The prompt generates signals, some of which become durable memory. On a new session, memory is pulled back into the prompt. Over time, retrieval patterns and formats the model should know reflexively move from memory into fine-tuning. Once fine-tuned, what’s worth retrieving changes: if the model knows a format reflexively, you no longer retrieve examples of that format each call. The agent improves by doing its job, and the loop closes.
Srivastava is blunt about the hierarchy of difficulty in building this. “The model is the easy part,” he said. “What you build around the model, the harness around the model that helps you store the right information at the right place and circulate among them, is the key architecture that you got to build.”
His point on reasoning is equally direct. If a model cannot reason over two, three, or five documents, it will not reason over fifty. More context does not help a model that cannot reason. Graph-based RAG can assist reasoning, but that is where reasoning support stops. The expectation that adding retrieval will somehow confer reasoning ability is, in his framing, a category error.
The framework’s strength is that it converts a vague argument about AI architecture into a per-artifact checklist. Its unresolved tension is that the checklist depends on a judgment — has this stopped changing? — that teams are bad at making in advance. Srivastava’s own examples show the failure mode is rarely a dramatic wrong call; it is a sequence of locally reasonable ones, each made by a different owner, none labeled as architecture. The competitive implications extend beyond any single deployment. Two things worth tracking are whether drift monitoring on fine-tuned “stable centers” becomes standard practice in claims and moderation pipelines, and whether access-scoped retrieval metadata — repo, user, permission — gets treated as a first-class design requirement rather than an afterthought. The task for AI teams is no longer just choosing a model; it is deciding, deliberately and repeatedly, where each piece of knowledge lives, and who owns the answer.
References:
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.
The news and data on this website are for reference only and do not constitute investment advice or an offer to buy or sell. Information is sourced from exchanges and public sources, and may be delayed, interrupted, or updated. While we strive for accuracy, we do not guarantee timeliness, correctness, or completeness. Content may include external links for which we are not responsible. By using this site, you agree that we and our partners are not liable for any losses. Investment carries full responsibility; please carefully assess risks and consult professionals. If there are errors in the content, please contact us for correction.

source
This is a newsfeed from leading technology publications. No additional editorial review has been performed before posting.

Leave a Reply