IT leaders grappling with the complex challenges of AI transformation must quickly ensure they have a clear plan for the role of cloud.

Off-prem approaches have been adopted as large language models (LLMs) demand expensive, capital investment. The cloud has allowed businesses to test out AI quickly, while keeping costs under control.

According to Foundry’s 2025 AI Priorities Study, 53% of companies plan to increase AI spending.1 But, as AI matures, how organisations invest in AI is changing.

“AI is not only saving time and money, but it’s also transforming operations for the better because it’s learning and it’s evolving,” says Mahesh Subramony, senior fellow design engineer, at AMD.

But businesses increasingly understand that not all AI workloads are suitable for the cloud. Security, privacy and regulation, as well as data sovereignty and application performance factors, require careful strategic consideration.

For most enterprises, a hybrid approach will be more effective.

Developments in AI reflect this. Although ever-larger LLMs, with several hundreds of billions of parameters, offer the prospect of greater precision, AI developers are also working to fine tune LLMs, or even shrink their models down, to deliver better performance on less costly hardware.

“Almost all of the generative AI experiences you see today are on the cloud and for good reason,” says Subramony.

“There’s a lot of work that needs to be done around model optimisation, newer models and better capabilities on the edge device. We’re working on that to help ensure we can deliver that, when you want lower latency, security, and less reliance on connectivity.”

Small language models can cost less to run, and increasingly compete with LLMs, especially for clearly defined tasks or business processes.

And machine learning powered by mixture of experts (MoE) bring together multiple, specialist models to take on a complex task, with lower overheads than generative AI (genAI).

Other techniques, such as fine-tuning models by quantization and model compression are producing smaller, more agile AI tools. This gives enterprises a new set of options, when it comes to deploying AI to solve their business challenges.

“These emerging trends in the models are very encouraging,” says Subramony. “They’re becoming more capable, with increased efficiency.”

Cloud alternatives, local choices

These developments in AI models are unlocking the potential for new deployment models: on-premises, on servers and PCs, and ultimately, on personal devices.

This allows organisations to overcome some of the very real security and privacy concerns. But there are performance and productivity benefits too.

The arrival of power-efficient AI processing is also opening up new options. Dedicated neural processing units, or NPUs, allow devices, such as laptops, to undertake a far wider range of AI tasks locally.

AI-capable chips from AMD, for example, support 40 plus TOPS (trillions of operations per second) from the NPU alone. And ever more power-efficient NPUs will support a wider set of AI use cases on local devices without impacting battery life: a critical issue for edge users.

“That’s the promise of the AI PC,” says Subramony.

“As it is embedded in devices and applications used across the enterprise, it’s going to be a powerful tool that helps organisations work faster, think smarter and deliver more impact.”

And the performance of AI PCs, driven by NPUs, is set to significantly improve over the next few years.

This stands to make local access to AI all but ubiquitous, even as AI OCs work together with the cloud for more demanding workloads.


1 Foundry, AI Priorities Survey 2025, https://foundryco.com/tools-for-marketers/research-ai-priorities/


Share
Share