Nvidia inference computing is entering a critical transition period as the artificial intelligence industry moves away from model training and toward running models at scale. The shift is reshaping competitive dynamics across the chip industry and raising questions about Nvidia’s long-term market position.
The Rise of Inference in AI Workloads
The artificial intelligence industry has moved into a new phase where running models — known as inference — is growing faster than training them. Nvidia Chief Executive Jensen Huang said at a March 4 investor conference that the shift toward agentic AI, which relies almost entirely on inference computing, has been visible for some time. Companies like OpenAI and Anthropic now produce thousands of times more inference tokens than before, Huang said.
Why Nvidia’s Current Chips Face Limitations
Nvidia’s Grace Blackwell servers consume large amounts of energy and lack sufficient memory for efficient inference workloads, according to users. Paul Kedrosky, a venture investor and research fellow at the Massachusetts Institute of Technology’s Initiative on the Digital Economy, said Nvidia is in a difficult position. He argued that Nvidia’s gross margins, which stood at 73% in the most recent quarter, will face downward pressure as inference computing prioritizes efficiency over raw performance.
“Nvidia is in a weird moment. For a long time, Jensen was saying, ‘We don’t need to have dedicated, stand-alone inference chips, you can just throw a Blackwell at it.’ But that ship has sailed, and there’s a host of new competitors.”
Paul Kedrosky, Venture Investor and Research Fellow, MIT Initiative on the Digital Economy
Nvidia’s $20 Billion Groq Deal and Strategic Pivot
In December, Nvidia paid $20 billion to license chip technology and hire talent from Groq, a startup that designs language processing units suited to running AI models. At its GTC conference this week, Nvidia plans to unveil a computing platform combining a modified Rubin GPU with a Groq processor tailored to inference computing. The deal was fast-tracked after OpenAI signed a $10 billion agreement with rival chip startup Cerebras.
In February, Meta Platforms said it would install thousands of Nvidia’s Vera CPUs in its AI data centers — the first significant deployment of Nvidia systems for AI that did not include GPUs. Nvidia is also planning additional computing solutions involving multiple CPUs unattached to GPUs, the Wall Street Journal reported. Meanwhile, Intel is teasing a major partnership announcement with Nvidia at the GTC event.
Competitors Encroach on Nvidia’s Market Share
Cerebras CEO Andrew Feldman has argued in public posts that Nvidia’s proprietary programming library CUDA is generally only needed for training, not inference. Last week, Cerebras signed Amazon Web Services as a new customer, further expanding its footprint in the cloud computing market. Tom Burke, chief revenue officer of U.K.-based cloud provider Nscale, said the compute mix has shifted dramatically.
“If you were to look at this market 12 months ago, it was probably 90-10 training versus inference in terms of the compute people needed. I think by the end of this year, it will have swung.”
Tom Burke, Chief Revenue Officer, Nscale
Shahriar Rabii, a former Google and Meta executive and co-founder of chip startup Majestic Labs, said the best quality models are increasingly not viable with existing infrastructure. The comment reflects a broader industry view that current hardware architectures require rethinking to meet inference demands efficiently.
Nvidia’s Leadership Remains Confident
Despite competitive pressure, Nvidia’s Chief Financial Officer Colette Kress said agentic AI workloads are becoming a major driver of revenue growth. She stated that Nvidia’s chips will continue to dominate for the foreseeable future. Huang reinforced that position during Nvidia’s most recent earnings call, stating that inference now directly translates into customer revenues as agents generate large volumes of tokens.
“Right now, we’re the king of inference.”
Colette Kress, Chief Financial Officer, Nvidia
How far ahead Nvidia remains in the AI infrastructure race depends largely on whether its new chips built with Groq prove fast, efficient, and affordable enough to outpace rivals. The outcome will determine whether Nvidia can successfully extend its dominance from the training era into the inference era.

