About cross-cloud data access

The cross-cloud data access feature lets you query data stored in other cloud providers directly from Google Cloud without migrating files or building complex ETL pipelines with Cross-Cloud Interconnect.

As part of borderless Lakehouse, this capability lets you perform unified analytics and apply AI across your distributed datasets using BigQuery, standalone Apache Spark environments, or Managed Service for Apache Spark.

In addition to analytical queries, you can use your federated data for AI-driven insights and governance:

  • Conversational Analytics: Build specialized agents grounded in your exact data sources, including cross-cloud tables, to analyze data across clouds from a single conversation.
  • Knowledge Catalog: Use Knowledge Catalog features for data profiling and insights with federated data sources.

Use cases

Lakehouse supports several key use cases for accessing data across multiple cloud providers:

  • Reduced data movement lets you query data stored in other cloud environments directly, simplifying data access and processing.
  • Unified analytics lets you perform advanced analytics with consistent features and hardware optimization across all your data, regardless of where it resides.
  • Borderless AI and ML lets you apply AI models, autonomous agents, and machine learning directly to your remote data without migrating it.

How accessing cross-cloud data works

Lakehouse queries remote data using the following process:

  1. Metadata discovery: Google Cloud's Lakehouse connects to remote Apache Iceberg REST catalogs, such as Databricks Unity or AWS Glue. Lakehouse discovers the data without copying any files. Depending on the remote catalog provider, Lakehouse authenticates securely through Secret Manager or OpenID Connect token federation with Google as the identity provider (OIDC token federation).
  2. Secure transport: Choosing to route traffic over a private interconnect (for example, Dedicated CCI or Partner Interconnect) significantly reduces data transfer costs compared to the public internet and makes latency highly predictable.
  3. Optimized execution: As queries read data from remote clouds, Lakehouse temporarily caches those data segments locally within Google Cloud on specialized storage. Subsequent queries use the local cache, which avoids a significant portion of cross-cloud egress charges.

Supported catalogs

Lakehouse supports querying data from the following remote catalog providers: