The AI race is shifting focus from building the largest models to developing smarter, more cost-effective systems. Companies like Perplexity are creating orchestration layers that intelligently select the best model for each task, rather than relying on a single, monolithic AI.
This trend towards efficient, adaptable AI systems is being driven by the increasing capabilities and lower costs of open-weight models, challenging the economic models of leading AI providers and potentially reshaping the future of AI deployment.
AI Race Shifts: Focus Moves from Bigger Models to Smarter, Cheaper Systems
The artificial intelligence landscape is evolving beyond a simple competition for the largest and most advanced models. Companies are now prioritizing the development of more efficient, cost-effective, and adaptable AI systems that can intelligently select the best model for specific tasks.
For the past two years, the artificial intelligence race has been characterized by a focus on developing bigger models and achieving higher benchmark scores. However, as AI transitions from experimental phases to real-world applications and workflows, the emphasis is shifting. The key is no longer just accessing the best model, but utilizing the one that is the most suitable for a particular job, considering factors like cost, data requirements, and operational environment.
This paradigm shift is fostering a new era of AI competition, one that is less about model size and more about efficient routing, cost optimization, robust control, and optimized compute usage.
"The model alone is no longer the product. It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools."
This means AI products are increasingly becoming sophisticated systems capable of discerning which model to use, when to deploy it, and what external tools or internal data sources are necessary. For instance, a customer service inquiry might not require the most expensive model, while a complex coding problem might. Routine internal tasks could be handled by more affordable open-weight models, with more demanding steps escalated to more powerful ones.
"The answer is always use whatever is the best for the task," Srinivas added.
The rise of alternative models coincides with corporations re-evaluating their AI spending. This presents a significant challenge for leading AI providers like OpenAI and Anthropic, who have previously thrived by offering the most advanced proprietary technology.
Perplexity recently unveiled a new system for its computer-use product, built around GLM 5.2, an open-weight model from China's Z.ai. This system is designed to maximize the use of less expensive models while only engaging more powerful ones when absolutely necessary.
This strategy aligns with a broader market trend: open-weight models, which can be freely downloaded, customized, and operated by companies, are rapidly increasing in capability and becoming more cost-effective compared to premium proprietary models.
Peter Fenton, a general partner at Benchmark, believes this shift could be transformative. "A maybe contrarian view that is becoming consensus is our belief that 90-plus percent of the tokens created will come out of open-weight models over the next 18 to 24 months, possibly even by the end of the year," Fenton stated. Tokens are the fundamental units of data that AI models process and generate.
Fenton further elaborated, "The inference margins generated by the frontier model companies, I think, are going to come under pressure when you can run those without the markup that they're providing, when you have good enough models from open weights." He also noted that smaller, task-specific models can sometimes outperform larger, general-purpose models in speed and performance.
Where and How AI Runs Matters
This focus on deployment and efficiency is a key reason why Benchmark invested in Ollama, a company that simplifies the process of downloading, running, and managing open-weight models for developers and enterprises.
Jeff Morgan, CEO of Ollama, highlighted the shift in priorities: "One thing is where the model's from and where it was created and trained. But the more important thing to these businesses we speak to is where it runs and how it runs."
Ollama has reportedly seen adoption by over 85% of Fortune 500 companies, including those in highly regulated sectors like aviation, insurance, and healthcare. Many organizations begin with smaller models operating within their own data environments, gradually scaling to larger open-weight models as their confidence and needs grow.
The growing prominence of open-weight models also presents strategic considerations for the U.S., as many leading open-source models originate from Chinese labs such as Z.ai and DeepSeek. This development elevates open-source AI to a critical business, policy, and national competitiveness issue.
Srinivas advocates for U.S. support of open models, emphasizing their role in making AI more affordable and accessible. "If you want the benefits of AI to be widely distributed to small businesses in America and American allied countries, then you really need AI to be a lot more affordable. And open source is the only way to do that," he stated.
This trend could also impact the massive data center buildouts occurring throughout the tech industry. The current AI boom largely assumes continued demand for large cloud data centers equipped with high-end chips. However, Srinivas suggests that a portion of AI processing might eventually shift to local devices, whether owned by consumers or businesses, leading to a more hybrid AI system where routine tasks are handled locally, and complex computations are routed to powerful cloud-based models.
For investors, the critical question remains whether the major AI labs can sustain their premium pricing strategies as open-weight models improve and companies adopt more selective approaches to AI utilization.
WATCH:
