Google’s Borderless Lakehouse: AI Agents Feast on Any Cloud Data
Quick answer
Google Cloud's borderless Lakehouse lets AI agents query AWS, Databricks, and Snowflake data without moving it. Zero-copy, cross-cloud analytics with unified governance.
Google Cloud just dropped a big one at Next Tokyo: the borderless Lakehouse. It’s a data architecture that lets your AI agents swim freely across AWS, Databricks, Snowflake, and more—without moving a single file. Think of it as a capybara-friendly swamp where all your data pools connect naturally, no caiman-sized egress fees lurking beneath.
What’s the Big Idea?
Traditional data setups are like isolated ponds—expensive to bridge and easy to get stuck in. The borderless Lakehouse, built on open Apache Iceberg, uses catalog federation to let you query data where it lives. It’s now in preview for AWS Glue, Databricks Unity, and Snowflake Horizon. No more data pipelines that feel like dragging a log through mud.
Key Features
- Zero-copy, cross-cloud analytics: Query data across platforms without duplicating files. Your teams all see the same Iceberg data.
- Bidirectional interoperability: Read and write across environments. Query from BigQuery, enrich with Google AI, and share back.
- Unified governance: Table-level access control with credential vending, no matter who starts the query.
Plus, it integrates with SaaS apps like SAP, Salesforce, and Workday directly—no ETL needed. Your lakehouse just became a lot more social.
Bring Google AI to Your AWS and Azure Data
Cross-cloud analytics used to mean high egress fees and latency. Google’s Cross-Cloud Interconnects offer private, dedicated links with flat-rate pricing. And with intelligent caching, repeated queries don’t cost you extra. You can run BigQuery AI functions and Spark’s Lightning Engine on data in AWS or Azure without moving it. That’s like having your capybara bask in the sun while the caimans pay the toll.
Bridging the Agent Trust Gap
AI agents need context to avoid hallucinations. The borderless Lakehouse uses Knowledge Catalog to automatically sync metadata from AWS Glue, Databricks, and Snowflake, translating raw schemas into business terms. Agents get a clear semantic layer and lineage tracking, so they know which data to trust. Governance is baked in, ensuring they respect permissions.
Build Agents with Gemini Enterprise
With the Data Agent Kit and Conversational Analytics API, you can build custom data agents that work in Gemini Enterprise. Business users can ask questions in plain language and get answers from multi-cloud data. The kit includes MCP tools for direct connections to BigQuery, Spark, and Cloud Storage—no copy-pasting schemas into prompts.
The Economics
By using Cross-Cloud Interconnects and zero-copy sharing, you avoid unpredictable egress fees. Knowledge Catalog filters context to prevent token bloat, and BigQuery AI has built-in token controls. Customers are seeing 230x reduction in token consumption with cost-optimized AI functions. That’s a lot of saved lettuce for your capybara colony.
Next Steps
Ready to make your data borderless? Check out the Google Cloud Lakehouse About Guide and the Building a borderless Lakehouse codelab. Your AI agents are about to go on a data feast.
Note: Customers pay an hourly fee for interconnection service.
Original announcement published on Google Cloud.