RAG Was Only the Beginning
The New AI Learning Stack Every Digital Leader Needs to Understand
CDO TIMES PEOPLE | DATA | AI | HIGHER IMPACT | EXECUTIVE |
Why inference-time learning, world models and the new AI learning stack change enterprise strategy
The next competitive advantage in enterprise AI will not come from choosing one “best” model. It will come from understanding where intelligence should adapt: in external context, in memory, in temporary model state, in persistent weights, or inside a simulated representation of the world.
BY CARSTEN KRAUSE | CDO TIMES | EXECUTIVE ANALYSIS |

FIGURE 1 | The enterprise AI adaptation stack: different mechanisms change different parts of the system.
What digital leaders should know now
|
The Architecture Question Is No Longer “Which Model?”
For the last several years, many enterprise AI strategies have effectively been reduced to a familiar formula: select a foundation model, connect it to corporate information through retrieval-augmented generation, add guardrails, and call the result an enterprise AI platform. That architecture remains useful, but it is no longer sufficient to describe what modern AI systems are becoming. RAG can give a model access to current information, but it normally does not teach the model anything. A reasoning model can spend more compute on a difficult problem without permanently changing what it knows. Meanwhile, test-time training, persistent agent memory, continual learning and world models are moving AI toward systems that can adapt, accumulate experience, specialize and simulate consequences in ways that traditional RAG architectures were never designed to support.
That distinction matters because digital leaders are about to face a more consequential architecture question than which LLM provider should sit behind an API. The better question is: where should learning and adaptation happen in our AI system, how long should the adaptation persist, and who should govern it? Once that question is asked, RAG, fine-tuning, inference-time reasoning, memory, continual learning and world models stop looking like competing buzzwords. They become separate mechanisms in an enterprise intelligence architecture, each with different economics, risks and operating models. Organizations that understand those differences will be better positioned to build AI systems that do more than retrieve yesterday’s knowledge and generate plausible answers.
THE DECISION RULE Keep volatile knowledge external. Make durable behavior internal only when it earns the right to persist. Treat adaptive model state as governed state. Simulate before acting when consequences matter. |
RAG Is Valuable: But RAG Is Not Learning
Retrieval-Augmented Generation remains one of the most practical techniques available to enterprises because it lets a model use information that was not embedded in its original training. The original RAG work combined a pretrained generative model with external, non-parametric memory, allowing the system to retrieve information and use it while generating an answer. In a modern enterprise implementation, that usually means searching documents, databases, policies, product information or other governed sources and inserting the relevant material into the model’s context. The crucial point is that the underlying model normally does not change as a result. The knowledge is accessed at runtime, not learned into the model’s durable parameters.
SOURCE Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks https://arxiv.org/abs/2005.11401 |
That difference is easy to overlook because the user experience can feel like learning. Ask an enterprise assistant about a policy uploaded ten minutes ago and it may answer correctly even though the foundation model was trained long before the policy existed. What happened was not retraining; information was found and placed into context. This is precisely why RAG remains attractive for regulated environments: knowledge can be updated, permissioned, traced and removed without rebuilding a model. It also gives security and compliance teams a clearer control surface because enterprise information can remain outside the model and be released only when the user, agent and task are authorized to see it.
Basic vector retrieval, however, is already evolving. Anthropic reported that its Contextual Retrieval approach reduced top-20 retrieval failures from 5.7% to 2.9% when contextual embeddings and contextual BM25 were combined, and to 1.9% when reranking was added — a 67% reduction versus its tested baseline. Microsoft’s GraphRAG addresses a different limitation by building a graph of entities, relationships and communities so systems can answer global questions that ordinary chunk retrieval often handles poorly. In one Microsoft Research evaluation on 50 AP News questions, dynamic community selection reduced token cost by an average of 77% at the tested community level while producing statistically similar quality to the static baseline. Those are meaningful architecture improvements, but they improve retrieval and context construction; they still should not be confused with the model itself learning the enterprise.
SOURCE Anthropic, Contextual Retrieval https://www.anthropic.com/engineering/contextual-retrieval |
SOURCE Microsoft Research, GraphRAG: Improving global search via dynamic community selection https://www.microsoft.com/en-us/research/blog/graphrag-improving-global-search-via-dynamic-community-selection/ |
“Inference Learning” Actually Covers Three Different Ideas
The phrase inference-time learning is increasingly used as if it described a single new capability, but leaders should separate at least three mechanisms. The first is in-context learning, where examples, instructions and information in the prompt change the model’s behavior during the current interaction without changing durable weights. The second is inference-time or test-time compute, where the system spends more computational effort exploring, reasoning, checking or revising before it commits to an answer. The third is test-time training or adaptation, where part of the model or a dedicated fast-weight component is actually updated during inference. Only that third category is learning in the conventional parameter-update sense, and even there the state may be intentionally temporary rather than permanent.
The distinction is not academic because the cost and governance profiles are very different. OpenAI reported with its o1 research that performance improved with more test-time compute as the model spent more time reasoning, demonstrating that additional inference work can materially change answer quality without requiring a new pretrained model. But more reasoning is not universally better: a 2025 Anthropic research collaboration constructed tasks where extending reasoning length reduced accuracy and identified failure modes such as distraction, overfitting to problem framing and difficulty maintaining focus. The executive takeaway is that inference-time compute should be treated as a controllable resource with task-specific evaluation, not as an automatic intelligence dial. Enterprises will need routing policies that decide when a problem warrants additional compute and when it simply adds cost or new failure modes.
SOURCE OpenAI, Learning to reason with LLMs https://openai.com/index/learning-to-reason-with-llms/ |
SOURCE Anthropic Alignment Science, Inverse Scaling in Test-Time Compute https://alignment.anthropic.com/2025/inverse-scaling/ |

FIGURE 2 | A practical primer for leaders: inference, RAG, world models, fine-tuning and reinforcement learning solve different problems.
Test-Time Training Blurs the Line Between Using and Training a Model
Test-time training goes further than giving the model more context or more reasoning tokens. The ICLR 2026 paper Test-Time Training Done Right describes systems that adapt part of the model’s weights — often called fast weights — during inference so they can store temporary memories of the current sequence. Its authors frame the technique as an alternative way to model long-context dependencies, with a learned fast-weight state updated while the system is operating. Separate research on test-time training for in-context learning has shown how explicit weight updates on examples supplied at test time can reduce the amount of data needed for some downstream tasks. This remains an emerging research direction rather than a default enterprise design pattern, but it introduces an important possibility: future AI systems may adapt internally to a task while they are executing it, without waiting for a conventional retraining cycle.
SOURCE ICLR 2026, Test-Time Training Done Right https://proceedings.iclr.cc/paper_files/paper/2026/hash/ffbfff81fb78e4bb558273b91e12e318-Abstract-Conference.html |
SOURCE Gozeten et al., Test-Time Training Provably Improves Transformers as In-context Learners https://arxiv.org/abs/2503.11842 |
For enterprises, that creates a governance problem that RAG alone does not have. If a system can change internal state during execution, leaders need to know what changed, which data caused the change, how long the change persists, whether it can leak across users or security boundaries, and how it can be reset or rolled back. RAG governance is largely about access control, retrieval quality, provenance and prompt construction. Test-time learning introduces model-state governance, which requires observability into an adaptive component that may not be visible through ordinary application logs. As these techniques mature, architecture teams should expect new requirements for state isolation, lineage, validation, rollback and policy enforcement at inference time.
Fine-Tuning Is Not Dead — It Is Becoming More Selective
RAG’s popularity led some organizations to treat fine-tuning as yesterday’s architecture. That is a mistake because retrieval and fine-tuning solve different problems. RAG is strongest when knowledge changes frequently or provenance matters, while fine-tuning is useful when the organization wants durable changes in behavior, terminology, classification, output structure or specialized task performance. Parameter-efficient approaches such as LoRA make this less expensive by freezing the base model and training comparatively small low-rank components instead of updating every parameter. The original LoRA paper reported dramatic reductions in trainable parameters and GPU memory for its tested models, reinforcing the broader point that specialization does not always require a full retraining program.
SOURCE Hu et al., LoRA: Low-Rank Adaptation of Large Language Models https://arxiv.org/abs/2106.09685 |
The more important strategic change may be in the training data itself. Google Research described an active-learning approach for ads safety that reduced its experimental training set from roughly 100,000 examples to fewer than 500 while increasing alignment with human experts by as much as 65%; Google also reported that larger production systems achieved reductions of up to four orders of magnitude while maintaining or improving quality. Those results should not be generalized into a universal promise because they depend on the domain, model and label quality. They do, however, reinforce a powerful enterprise principle: the most valuable proprietary training asset may not be a giant lake of undifferentiated content. It may be a much smaller set of high-information examples where expert judgment, edge cases and corrections materially change the desired outcome.
SOURCE Google Research, Achieving 10,000x training data reduction with high-fidelity labels https://research.google/blog/achieving-10000x-training-data-reduction-with-high-fidelity-labels/ |
Reinforcement Learning Adds Feedback, Not Just More Data
Reinforcement learning is another mechanism leaders should distinguish from retrieval and supervised fine-tuning because it changes behavior through feedback about outcomes rather than by simply showing the model more correct examples. Reinforcement Learning from Human Feedback became prominent through work such as InstructGPT, where human demonstrations and preference rankings were used to fine-tune a language model toward outputs people preferred. In that study, human evaluators preferred outputs from a 1.3-billion-parameter InstructGPT model to outputs from the much larger 175-billion-parameter GPT-3 on the researchers’ prompt distribution. The enterprise lesson is not that smaller models always beat larger ones; it is that training objectives and feedback signals can matter as much as raw scale. As agents increasingly take actions in workflows, reinforcement signals can come from human approvals, task success, policy checks, simulator results or downstream business outcomes — which makes reward design a governance issue, not just a machine-learning choice.
SOURCE Ouyang et al., Training language models to follow instructions with human feedback https://arxiv.org/abs/2203.02155 |
Continual Learning Is the Long-Term Prize
and a Governance Challenge
Traditional enterprise model development is episodic: train or procure a model, release it, then replace or fine-tune it later. Continual learning aims for something closer to how organizations actually operate, with systems acquiring knowledge and skills over time while preserving what they learned before. The fundamental challenge is catastrophic forgetting, where adaptation to new tasks can degrade performance on older capabilities. Surveys of continual learning for large language models describe approaches spanning continual pretraining, domain adaptation and continual fine-tuning, but the field remains technically difficult. That difficulty is one reason persistent external memory is often a safer near-term architecture than simply allowing production models to update themselves continuously.
SOURCE Shi et al., Continual Learning of Large Language Models: A Comprehensive Survey https://arxiv.org/abs/2404.16789 |
Google Research’s Nested Learning work offers a more recent view of the problem by describing a model as interconnected optimization processes operating at different levels and update rates. The stated goal is to improve continual learning and reduce catastrophic forgetting by treating architecture and optimization as part of a nested system rather than as entirely separate concerns. This is research, not a ready-made enterprise product pattern, but it points toward models that could someday learn on multiple timescales instead of relying on one static set of weights plus a context window. For executives, the harder question may become organizational rather than technical: if the enterprise AI can learn continuously, who is authorized to teach it, which experiences deserve to become persistent behavior, and what happens when the organization itself has inconsistent or low-quality decision practices?
SOURCE Google Research, Introducing Nested Learning: A new ML paradigm for continual learning https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/ |
Agent Memory Creates a Different Kind of Learning
Agents introduce another layer that does not fit neatly into either RAG or model training: persistent memory. An agent can retain facts about previous interactions, successful workflows, tool outputs, user preferences, decisions and completed tasks without changing the foundation model’s parameters. From the user’s perspective, the system has learned because it behaves differently based on experience. Technically, much of that learning may be happening in an external state store that can be inspected, permissioned and deleted. This makes memory a powerful enterprise control point because the organization can decide which experiences are stored, how long they persist, which agents may access them, and whether sensitive information can ever be promoted from short-term working context into long-term memory.
A mature agent architecture will therefore need more than one kind of memory. Working memory holds information needed for the current task; episodic memory captures prior experiences and outcomes; semantic memory stores reusable facts and relationships; and procedural memory can encode how successful work is performed. These mechanisms can create systems that become more useful with experience even when the foundation model remains unchanged. That distinction is strategically important because enterprises may be able to gain many of the benefits people casually describe as ‘continuous learning’ through governed external memory before they allow models to change their own weights. In many high-risk environments, inspectable memory will be a more defensible first step than invisible adaptation.
Distillation Turns Expensive Intelligence Into Specialized Capability
Not every task should be handled by the largest available frontier model. Once an enterprise identifies a repeatable task with stable requirements, knowledge distillation can transfer behavior from a larger teacher model into a smaller specialist model. Google Research’s Distilling Step-by-Step work reported that a 770-million-parameter T5 model outperformed a few-shot-prompted 540-billion-parameter PaLM model on the ANLI benchmark while using 80% of the available training data, a result tied to that benchmark rather than a general claim that small models outperform frontier systems. The architectural implication is more durable: a powerful general model can help create compact task-specific models that are cheaper to serve, easier to deploy close to the edge, and more predictable for narrow workloads. That makes distillation particularly relevant for enterprises facing large transaction volumes, latency constraints, privacy boundaries or infrastructure cost pressure.
SOURCE Google Research, Distilling step-by-step https://research.google/blog/distilling-step-by-step-outperforming-larger-language-models-with-less-training-data-and-smaller-model-sizes/ |
World Models Move AI From Retrieving Facts to Predicting Consequences
RAG and language models are primarily information-centric: they retrieve, transform and reason over representations of what has been written or encoded. World models attempt something more ambitious by learning a representation of how an environment behaves so a system can predict what may happen next and how actions could change the outcome. Google DeepMind’s Genie 3 demonstrated interactive generated environments running at 24 frames per second at 720p while maintaining consistency for several minutes. Meta’s V-JEPA 2 work took another path by learning from video and then using less than 62 hours of robot data from the DROID dataset to support zero-shot planning on robot arms in new environments. These are not proof that general enterprise world models are ready for every business process, but they show why simulation, prediction and planning are becoming distinct AI capabilities rather than extensions of document retrieval.
SOURCE Google DeepMind, Genie 3: A new frontier for world models https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/ |
SOURCE Meta AI, V-JEPA 2 https://ai.meta.com/research/vjepa/ |
The category is also moving beyond academic demonstrations. On September 1, 2026, World Labs introduced Atlas, which it describes as an omni world model pretrained to operate across text, images, video and 3D and designed to generate, reconstruct and simulate spatial environments. Atlas is entering early access with select partners, so it should be treated as an emerging capability rather than a mature enterprise platform. The strategic direction, however, is clear: AI systems are being built not only to answer questions about the world but to maintain representations of environments that can be explored and simulated. For industrial companies, retailers, utilities, transportation providers and asset-intensive businesses, the convergence of world models with digital twins may become substantially more important than another marginal improvement in chatbot quality.
SOURCE World Labs, Atlas: A World Model for Spatial Intelligence https://www.worldlabs.ai/blog/atlas |

FIGURE 3 | Where learning happens determines persistence, control requirements and governance impact.
The Enterprise AI Architecture Is Becoming a Learning Portfolio
The most useful way to compare these approaches is to ask what actually changes when the system encounters new information. If only the runtime context changes, the adaptation is easy to reverse and usually easier to govern. If external memory changes, the system can accumulate experience without altering the base model, but access controls and memory hygiene become critical. If temporary model state changes, the enterprise must govern runtime adaptation and state isolation. If persistent weights change, the organization is creating durable behavior that requires stronger evaluation, versioning and rollback. And if the AI builds or uses a world model, leaders have to govern not only what the model says, but the assumptions inside the simulated environment on which consequential decisions may be based.
Approach | What changes | Persistence | Best enterprise use | Primary leadership concern |
|---|---|---|---|---|
RAG / Contextual RAG | Retrieved runtime context | Temporary | Current policies, documents, product knowledge, traceable answers | Permissions, provenance, retrieval quality |
GraphRAG | Structured relational context | Temporary | Portfolio questions, cross-document relationships, global queries | Graph quality, cost, stale relationships |
In-context learning | Prompt-conditioned behavior | Prompt/session | Examples, instructions, rapid task adaptation | Prompt quality, context contamination |
Inference-time compute | Reasoning/search effort | Temporary | Complex reasoning, verification, planning | Cost, latency, overthinking failure modes |
Test-time training | Fast weights / adaptive state | Task/session dependent | Distribution shifts, long-context adaptation | State lineage, isolation, rollback |
Fine-tuning / PEFT | Persistent weights or adapters | Persistent | Domain behavior, structure, classification, specialized tasks | Training data quality, regression risk |
Continual learning | Model knowledge across learning cycles | Persistent | Changing domains and long-lived adaptive systems | Catastrophic forgetting, who may teach the system |
Agent memory | External persistent state | Persistent / inspectable | Long-running workflows, personalization, experience reuse | Retention, privacy, memory poisoning |
Distillation | New smaller specialist model | Persistent | High-volume, stable, latency-sensitive workloads | Drift from teacher, maintenance overhead |
World models | Representation of environment dynamics | Persistent | Simulation, robotics, digital twins, planning | Model assumptions, simulation validity, consequence risk |
What I Would Build If I Were Designing the Enterprise AI Stack Now
I would not begin by building dozens of specialized models or by treating every knowledge problem as a vector-database problem. I would begin with a strong general-purpose reasoning model, governed enterprise retrieval, secure tool access, persistent but controlled memory, and an evaluation layer capable of measuring both answer quality and real business outcomes. I would use RAG where information changes frequently and provenance matters, GraphRAG where relationships across information domains matter, and fine-tuning only when repeated prompting cannot reliably produce the desired behavior. I would treat true test-time training and open-ended continual learning as emerging capabilities that deserve controlled experimentation before they are allowed into high-risk production workflows. The priority is not to maximize how much the system can learn, but to maximize how much useful adaptation can occur while the enterprise still understands and controls the consequences.
I would also identify use cases where the core problem is not knowledge retrieval but consequence prediction. Those are the candidates for simulation, digital twins and eventually world-model architectures. Physical operations, supply chains, energy systems, network operations, robotics and complex portfolio planning all contain state transitions that cannot be captured adequately by retrieving another PDF. The system has to represent what exists now, what could change, what an action would alter and how uncertainty propagates. That shift from retrieving facts to modeling consequences may become one of the most important architecture distinctions of the next phase of enterprise AI.
Finally, I would build a learning control plane around the environment. Every AI system should be able to expose what enterprise knowledge it retrieved, which memory it used, which tools it called, how much inference-time compute it consumed, whether any adaptive state changed, what feedback was captured and whether that feedback is eligible to influence future behavior. The more adaptive AI becomes, the less acceptable it is for learning to happen invisibly. Observability, provenance, evaluation and human decision rights should be designed into the learning architecture from the beginning rather than layered on after deployment. This is where human intelligence and AI become complementary: machines can adapt and simulate at scale, but people remain accountable for what the enterprise chooses to reinforce, remember and act upon.

FIGURE 4 | The enterprise intelligence flywheel: data and human expertise create value only when reasoning, action, evaluation, memory and selective learning are connected.
The Strategic Shift: From Model Selection to Learning-System Design
The AI market still spends enormous energy comparing model benchmarks, but enterprise differentiation is increasingly determined by what surrounds the model. Two companies can have access to the same foundation model and achieve very different results because one has better context, cleaner permissions, stronger tools, better memory, better evaluation and clearer human decision rights. A model benchmark tells leaders what a model may be capable of in a controlled evaluation; it does not tell them whether the enterprise has built a system that improves safely with experience. The architecture discussion therefore has to move up a level from choosing a model to designing an intelligence system. The durable competitive advantage will come from how effectively the enterprise combines human expertise, data, memory, reasoning, feedback and simulation into a controlled learning loop.
RAG answers the question, ‘What should the model know right now?’ Fine-tuning asks, ‘What behavior should become durable?’ Inference-time compute asks, ‘How much reasoning effort is justified for this problem?’ Test-time training asks, ‘Should the system adapt internally to this specific situation?’ Agent memory asks, ‘What should the system remember from experience?’ Continual learning asks, ‘What should the model itself keep learning over time?’ World models add the most consequential question: ‘What is likely to happen if we take this action?’ Digital leaders do not need to become machine-learning researchers, but they do need enough architectural literacy to know when those are different questions and when a vendor is collapsing them into one marketing term.
THE CDO TIMES BOTTOM LINE RAG is not the end state of enterprise AI architecture. The winning enterprise stack will deliberately decide what stays external and retrievable, what becomes memory, what receives more reasoning compute, what is allowed to adapt temporarily, what becomes durable model behavior, and where simulation is required before action. The strategic question is shifting from “Which AI model are we deploying?” to “How does our enterprise intelligence learn, remember, reason, simulate and improve — and who controls each of those mechanisms?” |
A Five-Point Agenda for Digital Leaders
1. Classify the problem before choosing the technique.
Separate knowledge freshness, behavioral consistency, reasoning difficulty, persistent memory and consequence prediction. They are different requirements and should not default to the same architecture.
2. Create an adaptation policy.
Define which information may remain in transient context, which experiences may enter long-term memory, which model components may be fine-tuned, and whether any runtime adaptation is permitted.
3. Build evaluation before self-improvement.
Do not let agents or models learn from outcomes that the enterprise cannot measure. Reward signals, human feedback and automated evaluations should be designed around business value and risk.
4. Treat simulation as a new enterprise capability.
Identify high-consequence decisions where digital twins or world models could test options before real-world execution. Start with narrow, measurable domains rather than attempting to simulate the enterprise at once.
5. Put human judgment around the learning loop.
The more persistent the adaptation, the stronger the human accountability should become. Human expertise should determine what deserves to become institutional machine memory and what should remain temporary.
Research and Source Links
Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
https://arxiv.org/abs/2005.11401
Anthropic — Contextual Retrieval
https://www.anthropic.com/engineering/contextual-retrieval
Microsoft Research — GraphRAG: Improving global search via dynamic community selection
https://www.microsoft.com/en-us/research/blog/graphrag-improving-global-search-via-dynamic-community-selection/
OpenAI — Learning to reason with LLMs
https://openai.com/index/learning-to-reason-with-llms/
Anthropic Alignment Science — Inverse Scaling in Test-Time Compute
https://alignment.anthropic.com/2025/inverse-scaling/
ICLR 2026 — Test-Time Training Done Right
https://proceedings.iclr.cc/paper_files/paper/2026/hash/ffbfff81fb78e4bb558273b91e12e318-Abstract-Conference.html
Gozeten et al. — Test-Time Training Provably Improves Transformers as In-context Learners
https://arxiv.org/abs/2503.11842
Hu et al. — LoRA: Low-Rank Adaptation of Large Language Models
https://arxiv.org/abs/2106.09685
Google Research — Achieving 10,000x training data reduction with high-fidelity labels
https://research.google/blog/achieving-10000x-training-data-reduction-with-high-fidelity-labels/
Ouyang et al. — Training language models to follow instructions with human feedback
https://arxiv.org/abs/2203.02155
Shi et al. — Continual Learning of Large Language Models: A Comprehensive Survey
https://arxiv.org/abs/2404.16789
Google Research — Introducing Nested Learning
https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/
Google Research — Distilling step-by-step
https://research.google/blog/distilling-step-by-step-outperforming-larger-language-models-with-less-training-data-and-smaller-model-sizes/
Google DeepMind — Genie 3
https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/
Meta AI — V-JEPA 2
https://ai.meta.com/research/vjepa/
World Labs — Atlas: A World Model for Spatial Intelligence
https://www.worldlabs.ai/blog/atlas
CDO TIMES | STRATEGY TODAY. A MORE INTELLIGENT TOMORROW. |


