RIYADH – Zyphra, an AI research and product company, announced on Tuesday the successful training of ZAYA1, a large-scale Mixture-of-Experts (MoE) foundation model. The project distinguishes itself as the first major model trained entirely using AMD’s hardware ecosystem, specifically the AMD Instinct™ MI300X GPUs and Pensando™ networking, rather than the industry-standard NVIDIA infrastructure.
The release marks a significant technical validation for AMD’s hardware in the high-performance computing market. For Saudi Arabia’s burgeoning digital economy, which relies heavily on importing computational power for Vision 2030 initiatives, this development suggests a viable alternative to existing GPU supply chain constraints.
Benchmarking Performance
According to the technical report released by Zyphra, the ZAYA1-Base model utilizes a total of 8.3 billion parameters, with 760 million active parameters.
In performance benchmarks cited by the company, the model reportedly outperforms Meta’s Llama-3-8B and the OLMoE model across reasoning, mathematics, and coding tasks. Zyphra also claims the model rivals the performance of Alibaba’s Qwen3-4B and Google’s Gemma3-12B.
Technical Specifications and Infrastructure
The training cluster was built in collaboration with IBM, utilizing IBM Cloud’s fabric and storage architecture combined with the AMD ROCm™ open software stack.
A key technical advantage highlighted during the development was the memory capacity of the AMD Instinct MI300X GPU. The hardware features 192 GB of high-bandwidth memory. Zyphra engineers stated that this capacity allowed them to avoid “tensor sharding” a complex process usually required to split models across multiple chips thereby simplifying the training stack.
“Efficiency has always been a core guiding principle at Zyphra,” said Krithik Puthalath, CEO of Zyphra. Puthalath added that the results highlight the utility of co-designing model architectures with specific silicon and systems.
Implications for AI Infrastructure
The project addressed common bottlenecks in Large Language Model (LLM) training. Zyphra reported that the optimized distributed I/O on the AMD platform resulted in model save times that were 10 times faster than standard configurations.
Emad Barsoum, corporate vice president of AI and engineering at AMD, said the milestone demonstrates the flexibility of AMD Instinct GPUs for training complex, large-scale models.
For regional data centers and institutions such as KAUST or SDAIA, the validation of the MI300X for training not just inference provides data points for diversifying hardware procurement. As the Kingdom localizes Artificial Intelligence capabilities, reducing reliance on a single hardware vendor could improve negotiation leverage and supply chain resilience.

