AI Inference

vLLM Adds DeepSeek V4.1 Flash Support for NVIDIA and AMD GPUs
vLLM has launched immediate support for DeepSeek V4.1 Flash on NVIDIA and AMD GPUs, bringing native vision and 1M-token context to GCC developers.

OpenAI Details Jalapeno Chip Performance Benchmarks Against Nvidia Hardware
OpenAI published benchmark data showing its custom Jalapeno chip outperformed Nvidia GB300 hardware in speed and energy efficiency tests across three large language models.

Qualcomm Acquires Modular to Build Silicon-Agnostic AI Software
Qualcomm acquires Modular for $3.92 billion to build a silicon-agnostic compute layer, aiming to challenge Nvidia’s AI dominance.
