Ollama cloud has received a hardware upgrade, with the platform now running on NVIDIA’s B300 data center processors to support Moonshot AI’s Kimi K2.5 and Z AI’s GLM-5 models. The update delivers faster throughput and lower latency across both models while maintaining reliable tool call support for third-party integrations.
Performance Gains from NVIDIA B300 Hardware
The switch to NVIDIA B300 processors directly improves response speeds for developers running inference workloads on the platform. Moreover, the upgrade preserves compatibility with existing tool call workflows, which are critical for production software integrations.
Ollama Cloud Integration with OpenClaw
Ollama has been officially integrated into OpenClaw’s default onboarding flow. Kimi K2.5 is listed as one of the recommended models for use with OpenClaw, and developers can launch it using the following command: ollama launch openclaw --model kimi-k2.5:cloud. By default, Ollama enhances OpenClaw with web search capabilities, providing access to current information during sessions.
Over 45,000 GitHub Integrations Supported
The Ollama cloud platform supports over 45,000 custom integrations sourced from GitHub, all accessible through the platform’s standard launch command. Consequently, developers working across a wide range of toolchains can connect to the upgraded models without modifying existing workflows. This broad compatibility positions the platform as a flexible backend for artificial intelligence development pipelines.
Subscription Tiers and Pricing Structure
Ollama cloud offers three fixed subscription tiers priced at $0, $20, and $100 per month. The company said the fixed-rate structure eliminates unexpected overage charges for users who leave tools such as Claude Code or OpenClaw running continuously. A free tier is available for users who want to evaluate the models before committing to a paid plan. Furthermore, Ollama stated that additional plans targeting higher usage volumes are in development.
“Ollama’s cloud comes with fixed subscription rates at $0, $20, and $100. This means you won’t wake up to surprise overage bills if you leave Claude Code or OpenClaw running.”
Ollama, Official Statement
Context and Market Position
Ollama is a platform that allows developers to run large language models locally and, through its cloud service, remotely at scale. The addition of NVIDIA B300 support reflects a broader industry trend of deploying the latest data center silicon to reduce inference costs and improve model responsiveness. Kimi K2.5, developed by Moonshot AI, and GLM-5, developed by Z AI, are both large-scale language models used in agentic and conversational AI applications. The hardware upgrade makes both models more accessible to developers building on the Ollama ecosystem.

