AMD has expanded its AI infrastructure portfolio with the launch of Helios, an open, rackscale AI infrastructure designed for frontier AI and sovereign computing. Helios is built around AMD's next-generation Instinct GPUs, EPYC Venice processors, Pensando networking, and the ROCm software stack. This marks AMD's first complete AI rack system with GPUs, CPUs, and networking built together, rather than selling separate chips. It is well suited for training large AI models, memory-heavy models, long context processing, and high-volume inference. According to Pareekh Jain, CEO at EIIRTrend & Pareekh Consulting, Helios is AMD's biggest shot yet at challenging Nvidia's dominance.
AMD has also secured an early hyperscale deployment for Helios, with Microsoft agreeing to deploy it to power its frontier model AI inference, its AI customers, and support Azure AI services. This strategic partnership underscores the growing demand for alternative AI infrastructure solutions beyond Nvidia's ecosystem.
The architecture behind Helios
The launch of Helios marks AMD's latest attempt to strengthen its position in a market where Nvidia continues to dominate AI infrastructure. Unlike previous AMD AI offerings centred on individual accelerators, Helios is designed as a complete rack-scale system integrating compute, networking, and software. The system includes 72 AMD Instinct MI455X GPUs with AMD EPYC Venice CPUs and AMD Pensando Vulcano networking using UALink, optimized for compute, data movement, and system efficiency. The platform also supports both OCP and MX data types, delivering up to 2.9 EFLOPS of FP4 and 1.4 EFLOPS of FP8 compute for AI training and inference. It integrates 31TB of HBM4 memory with 19.6TB/s of memory bandwidth, while a liquid-cooling design uses quick-disconnect connections to efficiently dissipate heat.
Helios is designed on open standards including OCP Open Rack Wide (ORW), Ultra Accelerator Link (UALink), and Ultra Ethernet Consortium (UEC) and can be scaled efficiently across datacenters while optimizing power, cooling, and serviceability for modern AI infrastructure. The use of open standards is a key differentiator from Nvidia's proprietary NVLink technology, giving buyers more flexibility in building heterogeneous data centers. According to Jain, Helios goes up against Nvidia's Vera Rubin rack. He notes that Nvidia is faster on raw inference speed and has a faster internal connection between chips, whereas AMD wins on memory size and offers better value for the price and power used. Helios's standout feature is memory, where each rack packs about 50% more total memory than Nvidia's competing system, which helps run very large AI models.
On the security front, Helios incorporates a hardware root of trust and continuous attestation at every layer. It supports hardware-enforced isolation, encrypted memory, and interconnects to help protect AI models, data, and workloads in multi-tenant environments. This is increasingly important as enterprises seek to deploy AI in shared cloud infrastructure while maintaining data sovereignty and compliance.
The software challenge
While the launch of Helios might help AMD close the hardware gap with Nvidia's rack-scale systems, software compatibility remains the real driver of enterprise adoption. AMD is expanding its ROCm AI software platform to support frameworks including PyTorch, TensorFlow, and JAX, enabling high-throughput inference and efficient distributed training while preserving familiar developer workflows. However, Jain emphasizes that while hardware parity or superiority in memory bandwidth is achievable, software maturity remains the key differentiator for Nvidia. Nvidia's CUDA software has a 15-20 year head start, and almost every AI tool, tutorial, and codebase defaults to it. He adds that software has been AMD's weak spot. AMD has improved ROCm a lot, but it still lags behind on the newest, most specialized optimizations, and setup is more complicated. For everyday AI work, ROCm is usable, but for cutting-edge performance, CUDA still leads.
The software ecosystem is critical for CIOs evaluating AI infrastructure. The availability of optimized libraries, pretrained models, and developer tools can significantly impact time to production. AMD has been actively working with the open-source community and major AI framework developers to improve ROCm compatibility. The company has also invested in tools like the AMD ROCm Software Development Kit and support for popular deep learning frameworks. However, the inertia behind CUDA remains formidable, and many enterprises will need to carefully assess whether their existing AI workloads can be migrated to AMD's platform without significant reengineering.
Evaluating the trade-offs
For CIOs evaluating AI infrastructure, Helios launch brings in another option to a market that has largely revolved around Nvidia's dominance. When considering Helios, CIOs will have to evaluate factors such as performance, software readiness, deployment models, procurement timelines, and total cost of ownership before committing to a platform. While AMD has not publicly announced a specific price tag for Helios, Jain believes it to be noticeably cheaper to buy and run, with lower chip prices and lower power use per GPU. This could be a significant factor as enterprises face skyrocketing AI infrastructure costs. The ability to achieve comparable performance at a lower cost could accelerate adoption, especially for organizations with budget constraints or those building large-scale AI clusters.
Helios gives companies a real second option besides Nvidia, easing supply shortages and providing leverage in negotiations. The catch is software, where teams need to check whether their AI tools run well on AMD's stack, since some advanced tools are still CUDA only. For CIOs planning to deploy both, Jain warns that the two systems cannot be plugged together into one combined machine as they use different, incompatible connection technology. However, companies can and do run both side by side in the same data center, just as separate systems handling different jobs. This hybrid approach allows enterprises to leverage the best of both worlds—using Nvidia for workloads that require the highest performance and CUDA compatibility, while reserving AMD for memory-intensive tasks or cost-sensitive deployments.
Another important consideration is the broader ecosystem support. AMD has been working with a range of OEMs and system integrators to offer Helios as part of validated solutions. This includes partnerships with Dell, HPE, Lenovo, and Supermicro, ensuring that enterprises have multiple procurement options. Additionally, AMD is investing in support for AI software stacks like PyTorch, TensorFlow, and JAX, with optimizations for distributed training and inference. The company has also launched the AMD ROCm Software Development Kit and is collaborating with the open-source community to address compatibility gaps. However, the maturity of these tools varies, and enterprises should conduct thorough testing before committing to large-scale deployments.
From a performance perspective, Helios's memory advantage is particularly relevant for large language models (LLMs) and foundation models that require massive memory capacity. With 31TB of HBM4 memory, each Helios rack can accommodate models with hundreds of billions of parameters, such as GPT-3 or LLaMA-scale models. The high memory bandwidth also benefits long-context processing, where models need to attend to large sequences of tokens. This positions Helios as a strong contender for frontier AI research and sovereign computing initiatives, where governments and organizations want to maintain control over their AI capabilities. The open standard design also facilitates integration with existing data center infrastructure, reducing vendor lock-in and enabling multi-vendor strategies.
The timing of the Helios launch is noteworthy, coming just as Nvidia faces potential delays in its next-generation Rubin architecture. This gives AMD a window to capture market share and demonstrate the viability of its platform. However, Nvidia's dominance extends beyond hardware to include a comprehensive software stack, developer tools, and a vast ecosystem of partners. AMD will need to continue investing in software and developer relations to close that gap. The company has already made significant progress with ROCm, including support for leading AI frameworks and the release of open-source libraries. But the ultimate test will be how quickly enterprises are willing to adopt AMD's platform for production AI workloads.
In conclusion, Helios represents a significant milestone for AMD in its quest to challenge Nvidia's AI infrastructure hegemony. The system offers compelling hardware advantages, particularly in memory capacity and cost efficiency, and Microsoft's early adoption provides a strong endorsement. However, the software ecosystem remains a critical hurdle that AMD must overcome to achieve widespread enterprise adoption. As CIOs evaluate their AI infrastructure strategies, Helios provides a credible alternative that can drive competition, innovation, and cost savings in the AI hardware market. The next few years will be decisive in determining whether AMD can translate its hardware prowess into lasting market share gains, especially as AI workloads continue to scale and demand increasingly specialized infrastructure.
Source: Network World News