Enterprise Ai After Hype
Captured source
source ↗Enterprise AI after the hype curve | AI21
Skip to Main Menu
Skip to Main Content
Skip to Footer
Back to Blog
-->
Back to Blog
A year ago, enterprise AI predictions focused on agents, RAG, multimodality, and continued gains from ever-larger models. Much of that direction held. What changed in 2025 was not what AI could do, but what determined whether it worked inside real organizations.
The year did not deliver a step change in top tier model capability. It delivered clarity about constraints. Top tier models remained strong, benchmarks stayed high, and performance gaps narrowed. In our experience, once these models hit real workflows, those gaps rarely decide outcomes on their own.
The year in numbers (symptoms, not explanations)
By the end of 2025, enterprise AI looked paradoxical on paper. Exploration was widespread. Agent adoption was actively discussed. And yet, very few organizations managed to scale AI systems inside core business functions.
These numbers don’t explain why this happened. They describe the symptoms of a gap between ambition and execution. Here’s what we learned about that gap.
Models converged, and benchmarks stopped helping
By the end of 2025, there was no significant improvement in top tier LLMs that translated into new enterprise outcomes. Benchmark results were impressive, but closely matched across leading models, and hard to translate into business impact.
Many companies we speak with understand that benchmarks do not mean much to their business as a decision tool. When results are this close, it is fair to ask whether benchmarks reflect enterprise reality, or whether they are being optimized toward in ways that do not map cleanly to real workloads.
So teams invested in what they cannot outsource, internal evaluation. They built proprietary test sets from their own workflows, created internal data assets, and used those to measure system performance end to end. In practice, this replaced picking a model based on a public leaderboard.
Reasoning reached limits, systems took over
We are still in the era of reasoning models, but we are transitioning to AI systems, because these systems can handle tasks a single LLM cannot reliably accomplish on its own. Across the teams we see making progress, the ingredients repeat.
Tools, persistent context, and retrieval. Multi-step execution with checks between steps. Better code generation. The shift is that teams stopped asking one model to do everything in one pass. They broke work into steps and gave the system a way to fetch context, call tools, validate intermediate results, and complete tasks end to end. In practice, that looks like investigating incidents across logs, resolving a customer case with policy grounding, or turning a spec into working code.
Enterprise adoption stayed internal, RAG first, and cautious on agents
Enterprises are still struggling with AI adoption. In 2025, most production usage concentrated on internal use cases built around a RAG pipeline, largely because reliability issues are more forgivable internally. The most prominent use cases were consistent, AI coding assistants, customer support, and internal Q&A chatbots.
We are not in the agentic era yet, at least not in enterprise. There was real momentum in pilots, and some teams scaled narrow agentic workflows, but broader autonomy remained limited in practice.
ROI pressure elevated SLMs
By late 2025, AI felt more mature, and ROI questions got louder. Efficiency moved from optimization to requirement. This is where the surge of interest in SLMs (Special or Small Language Models) makes sense. They are positioned as efficient, fast, and affordable.
You can see evidence in offerings like Thinking Machine’s Tinker, AWS’s Nova Forge, Jamba 3B, and many startups building similar approaches.
SLMs as enforcement layers in AI systems
SLMs also play an essential role inside AI systems. Many responsibilities benefit more from consistency and efficiency than open ended reasoning and higher costs. Use cases like guardrails, policy enforcement, security, and content moderation can be addressed with SLMs under current technology.
Looking forward SLMs will play an essential part in AI systems, and that role is expected to deepen into 2026 as systems develop further.
The open source ecosystem began to split
In 2025, open source AI started to diverge. Chinese AI models have been becoming more performant, including Kimi K2 , Qwen , and Deepseek . At the same time, western open models became more rare, a signal of a maturing field with sharper commercial incentives. As costs surged, doubts grew around whether open source model releases were ROI positive. Evidence for that is the fact that there is growing expectation that Meta may stop releasing Llama openly and focus instead on closed models, reportedly under the name avocado .
Dozens of agents create a need for orchestration
As multi-step systems grow, the need for an orchestrator becomes obvious. If you have dozens of agents in production, something has to break down complex tasks, route each part to a specialized agent (which can also be an SLM), and manage the context needed to complete the job.
At that point, orchestration stops being infrastructure and becomes the system. An orchestrator plus sub-agents is effectively an AI system, with coordination logic doing the work a single model cannot.
Data became the advantage, and model choice stayed unsolved
Organizations became more aware of the importance of their data in enabling AI applications. Data quality, availability, and uniqueness shape whether AI systems succeed. Organizational data contains business environment and context that shifts a generic model into something domestic and useful. This shows up in RAG pipelines and also in fine tuning.
As the field matures and efficiency constraints rise, model choice becomes a key trend. Choosing the right model for the right task under real constraints is still unsolved. Orchestration offers a solution by routing tasks to the right model at the right time.
What 2025 clarified
2025 did not crown a winner in models. But it did surface a few standouts, with Gemini 3 notably overperforming on several public evaluations. It clarified what blocks enterprise progress and what kinds of systems move forward. In our experience, the teams that made progress treated AI as a system, grounded in data, evaluated internally, and designed to behave consistently. With 2026 likely...
Excerpt shown — open the source for the full document.
Notability
notability 5.0/10Substantive blog post on enterprise AI.