Google Cloud announced a next-generation data lakehouse designed to support AI agents with real-time access to enterprise data across cloud boundaries. The platform combines fully managed Apache Iceberg storage with cross-cloud interoperability and AI-powered context capabilities.

The new architecture replaces traditional batch processing with continuous feedback loops and live data streams. This shift enables AI agents to access reliable context needed to transform raw data into actionable insights, unlocking both structured and unstructured enterprise data across multiple cloud environments.

Core Technical Capabilities

Google’s data lakehouse delivers four primary features. Fully managed Iceberg storage provides open-source flexibility combined with enterprise-grade performance, scale, governance, and multimodal processing capabilities. The platform introduces new cross-cloud interoperability, bringing Google’s high-performance foundation and AI capabilities to data stored on AWS and Azure.

A high-performance Apache Spark experience accelerates data science workloads with exceptional performance and developer environment flexibility. Additionally, AI-powered, always-on context enables AI agents to reason in real time across both operational and analytical data sources.

Cross-Cloud Integration and Interoperability

The platform addresses production cross-cloud data access challenges through new high-performance, scalable capabilities. Lakehouse cross-cloud interconnect and cross-cloud caching provide BigQuery and Managed Service for Apache Spark with high-performance access to AWS Iceberg data at scale, delivering price-performance characteristics similar to cloud-native solutions.

Lakehouse catalog federation for AWS Glue, Databricks, SAP, and Snowflake enables simple data discovery and analysis across any engine or cloud. The expanding partner ecosystem includes bi-directional access for Databricks, Oracle Autonomous Database, and Snowflake pipeline support for dbt, with Confluent Tableflow integration coming later in 2026.

Enterprise Adoption and Performance Metrics

Google stated the agentic-first lakehouse approach can deliver an estimated 117 percent return on investment with payback in under six months. Spotify is already leveraging Google Cloud’s Apache Iceberg products to build a modern data lakehouse that removes silos between data lakes and warehouses.

“Spotify is leveraging Google Cloud’s Apache Iceberg products as part of our efforts to build a truly modern data lakehouse that removes the silos between our data lakes and warehouses. This architecture provides us with an interoperable and abstracted storage interface, allowing our teams to process the same data across BigQuery, Dataflow, and other open-source engines without duplication.”

Ed Byne, Product Manager, Spotify

Accenture, a key partner in the initiative, views this as a fundamental shift in enterprise operations. The company stated that by utilizing Google Cloud’s lakehouse and zero-copy innovation, organizations can activate agentic AI with precision across industries including retail and life sciences.

AI-Powered Context and Agent Activation

Google Cloud’s Knowledge Catalog builds a unified foundation by aggregating business context from the entire data landscape, including Iceberg tables. The system delivers continuous enrichment by learning how enterprises actually use data, using Smart Storage to automatically map complex relationships within unstructured files.

BigQuery provides built-in Conversational Analytics, a Data Engineering Agent, and a Data Science Agent that work with cross-cloud, multimodal data. Looker Conversational Analytics delivers insights in natural language to business users. Organizations can build custom agents using Google-native tools like Agent Developer Kit and Model Context Protocol.

Real-time operational data integration is supported through Spanner, AlloyDB, and Cloud SQL, which enable real-time change replication into BigQuery and Iceberg. Analytical data in Iceberg can also be served with low latency using AlloyDB and Spanner.

Unified Governance and Multimodal Support

The platform introduces BigQuery ObjectRefs to merge unstructured data in Cloud Storage with structured data in Iceberg, simplifying multimodal analysis and managing conversational insights through BigQuery AI. Open lakehouse governance via Knowledge Catalog enhances enterprise trust through end-to-end data lineage, search, quality profiling, and table-level access controls for Iceberg estates.

Managed Service for Apache Spark offers a unified, high-performance experience accelerating data engineering to agentic AI development. Lightning Engine for Apache Spark delivers up to 2x the price-performance over leading high-speed Spark alternatives through vectorized execution, intelligent caching, and optimized I/O without requiring code changes.

Source: Google Cloud Blog