Large Language Models

ByteDance Develops Ten Trillion Parameter AI Model to Rival Mythos
ByteDance is reportedly training a new AI model with 10 trillion parameters, aiming to surpass Anthropic’s Mythos system and domestic Chinese competitors.

Kimi K3 Launched as the World’s First Open 3T-Class AI Model
Moonshot AI has released Kimi K3, a groundbreaking 2.8T-parameter open model featuring native vision and a 1-million-token context window.

New HalluSquatting Attack Exploits Popular AI Coding Tools to Build Botnets
Security researchers have uncovered a new vulnerability called HalluSquatting that allows attackers to hijack AI coding assistants and build massive botnets.

MiniMax M3 model Teased with 15.6X Decoding Speedup
MiniMax has teased its upcoming M3 model, featuring a custom sparse attention mechanism that delivers a 15.6x decoding speedup at a million tokens.

Google Introduces Gemini 3.5 Flash Model for Agents and Coding
Google unveiled Gemini 3.5 Flash, a model designed for coding and agents that matches leading competitors’ performance at lower cost. The model is now available in Google Search and the Gemini App.

Cursor Releases Composer 2.5, Doubling Usage for One Week
Cursor released Composer 2.5, an updated AI model offering 10x greater efficiency. The company doubles user access for one week and partners with SpaceXAI on a significantly larger model.

Gemini 3.1 Flash-Lite Now Generally Available on Enterprise Agent Platform
Google’s Gemini 3.1 Flash-Lite model is now generally available for enterprise use, designed for ultra-low latency and cost-efficient operations across software development, customer service, gaming, and financial applications.

xAI Releases Grok 4.3 with Always-On Reasoning and One Million Token Context
xAI shipped Grok 4.3 out of beta with one million token context window, always-on reasoning, and Custom Voices voice cloning suite. The company also launched Imagine Agent Mode directly on the web interface.

Alibaba Qwen Releases Qwen-Scope, Open Suite of Sparse Autoencoders
Alibaba releases Qwen-Scope, an open suite of sparse autoencoders enabling direct manipulation of model features, data classification, and model behavior analysis without prompt engineering.

Meta Launches Muse Spark, Its First Model From Meta Superintelligence Labs
Meta Superintelligence Labs launched Muse Spark on April 9, 2026, its first large language model built to power Meta AI across apps, glasses, and the web. The model supports multimodal perception, parallel agents, and a new shopping mode starting in the US.

US DoJ Appeals Court Ruling That Blocked Anthropic’s National Security Designation
The US Department of Justice appealed a court ruling that blocked the Trump administration’s national security designation against AI company Anthropic. The dispute stems from a $200 million contract cancellation over military AI deployment rights.

Marc Andreessen Says AI Is an 80-Year Overnight Success, Not Another Hype Cycle
Marc Andreessen joined the Latent Space podcast at a16z’s Sand Hill Road office to argue that AI represents decades of compounding progress, not a repeat hype cycle. He covered agents, open source, chip shortages, and why proving human identity online may soon require biometrics.

Qwen3.6-Plus Launches with 1M Context Window and Next-Level Agentic Coding
Alibaba’s Qwen3.6-Plus arrives with a 1M context window, top-tier agentic coding scores, and sharper multimodal vision. Open-source variants are coming soon.

Gemma 4: Google’s Most Capable Open AI Models Are Here
Google has launched Gemma 4, its most capable open model family yet, available in four sizes from mobile-optimized edge models to a 31B dense model ranking third globally. Released under Apache 2.0, it supports 140+ languages and advanced agentic workflows.

Xiaomi MiMo-V2-Pro: A 1-Trillion Parameter Model Challenging Western AI Leaders on Cost and Performance
Xiaomi released MiMo-V2-Pro, a 1-trillion parameter AI model, on March 18, 2026, ranking 10th globally on Artificial Analysis’s Intelligence Index at roughly one-seventh the cost of comparable Western models.

MiniMax M2.7 Model Launches on Agent Platform and API with Self-Improvement Capabilities
MiniMax has released its M2.7 model via the MiniMax Agent and API Platform, featuring reinforcement learning-based self-improvement and a 97% skill adherence rate across 40+ complex tasks.

Mistral Small 4 Unifies Reasoning, Multimodal, and Coding in One Open-Source Model
Mistral AI has released Mistral Small 4, an open-source model combining reasoning, multimodal, and coding capabilities under the Apache 2.0 license. The model delivers a 40% reduction in completion time and three times more throughput than its predecessor.

AI Token Pricing: How Context Became the Most Expensive Word in Enterprise AI
AI token pricing has made “context” the central billing unit in enterprise AI, with model output costs ranging from $2.50 to $180 per million tokens across major providers. The market prices capacity but not the quality of relational structure within those tokens.

LinkedIn Deploys LLM-Powered Ranking System to Personalize Feed for 1.3 Billion Users
LinkedIn has deployed a new LLM-powered Feed ranking system serving 1.3 billion professionals, replacing a fragmented multi-source architecture with unified retrieval and a sequential Generative Recommender model.

NVIDIA Releases Nemotron 3 Super: A 120B Hybrid MoE Model for Agentic AI Workloads
NVIDIA released Nemotron 3 Super on March 11, 2026, a 120B-parameter hybrid MoE model with a 1 million-token context window designed for multi-agent AI applications. The fully open model scores 85.6% on PinchBench and delivers over 5x the throughput of its predecessor.

OpenAI Releases GPT-5.4 With Native Computer-Use Capabilities and Improved Reasoning
OpenAI released GPT-5.4 on February 12, 2026, introducing native computer-use capabilities and a 1-million-token context window. The model achieves an 83.0% win rate against industry professionals on the GDPval benchmark.

Google Launches Gemini 3.1 Flash-Lite for High-Volume Developer Workloads
Google launched Gemini 3.1 Flash-Lite on March 3, 2026, offering developers a cost-efficient model priced at $0.25 per million input tokens with 2.5x faster response speeds than its predecessor.
