Skip to main content

Databricks Data Intelligence Platform on GCP

Overview​

The Databricks Data Intelligence Platform on Google Cloud provides a unified architecture for ingesting, transforming, governing, and serving data and AI workloads across the supply chain. It combines a lakehouse foundation (Delta Lake, Unity Catalog) with AI/ML capabilities (Mosaic AI), operational analytics, and a governed integration layer — all running on Google Cloud Storage as the persistence tier.

The architecture is structured across seven horizontal phases: Sources → Ingest → Transform → Query/Process → Serve → Analyse, with a vertical Integrate column for cross-cutting concerns (identity, governance, AI services, orchestration).

Architecture Diagram​

Sources
Ingest
Transform
Query / Process · Serve
Analyse
Storage
Integrate
ETL
Files / Logs
Sensors & IoT
RDBMS / DWH
Business Apps
Mediaunstructured
Federation
HMS*Hive Metastore
BigQuery
Sharing
Marketplaces
Data Shares
* external
Batch & Streaming — ETL
Lakeflow ConnectAuto Loader
Batch & Streaming — Federation
Cloud Data FusionPub/SubDatastream
↓
Databricks Data Intelligence Platform
Orchestration · CI/CD · MLOps · DataOps
Lakeflow JobsMLflowAsset BundlesSDKsTerraform Provider
Data Science, AI/ML & GenAI Apps — Mosaic AI
Feature EngineeringTraditional MLAgent BricksAgent FrameworkVector SearchModel ServingAI Gateway
AI / BI
DashboardsGenie
Data Engineering & Processing
PipelinesSpark / Photon
Data Warehousing
AI FunctionsDatabricks SQLConnectors & APIs
Data Intelligence
AssistantPredictive OptimizationPredictive IO
Operational DB
Cloud BigTableCloud SQLData Store
Applications
Databricks AppsBusiness / AI App
BI
Looker
Data Consumer
Sharing Partner
Data & AI Governance — Unity Catalog
FederationAccess ControlCatalog & LineageData & AI AssetsBusiness SemanticsQuality Monitoring
Data Management
Delta LakeIceberg
bronze→silver→gold
Collaboration
Delta SharingMarketplaceClean Rooms
Identity
ID Provider
Governance
Enterprise Catalog
AI Services
Anthropic
Vertex AI
Orchestration
External Orchestrator
Legend
Databricks
GCP / 3rd party
Google Cloud Storage (GCS) — unified object store for all Delta Lake, Iceberg, and raw data
Storage

Component Descriptions​

Sources​

SourceTypeIngestion path
Files / LogsSemi-structuredETL via Auto Loader or Lakeflow Connect
Sensors & IoTUnstructuredETL via Pub/Sub + Datastream
RDBMS / DWHStructuredETL via Lakeflow Connect / Cloud Data Fusion
Business AppsStructuredETL via Lakeflow Connect
MediaUnstructuredETL / Federation
HMS* / BigQueryFederated catalogFederation via Unity Catalog
Marketplaces / Data SharesExternal datasetsSharing via Delta Sharing

Ingest​

ComponentRole
Lakeflow ConnectNative managed connector for batch and CDC ingestion from databases and SaaS sources
Auto LoaderIncremental file-based ingestion from cloud storage with schema inference
Cloud Data FusionGCP-native ETL/ELT pipelines for complex data integration flows
Pub/SubGCP managed message bus for real-time event ingestion
DatastreamGCP CDC (change data capture) service for streaming database changes

Platform layers​

LayerKey componentsPurpose
OrchestrationLakeflow Jobs, MLflow, Asset Bundles, SDKs, Terraform ProviderCI/CD, job scheduling, MLOps lifecycle, IaC
Mosaic AIFeature Engineering, Traditional ML, Agent Bricks, Agent Framework, Vector Search, Model Serving, AI GatewayEnd-to-end AI/ML and GenAI development and serving
Data EngineeringPipelines, Spark / PhotonBatch and streaming data transformation at scale
Data WarehousingAI Functions, Databricks SQL, Connectors & APIsSQL analytics, embedded AI functions, external connectivity
Data IntelligenceAssistant, Predictive Optimization, Predictive IONatural language interface, automated performance tuning
Unity CatalogFederation, Access Control, Catalog & Lineage, Data & AI Assets, Business Semantics, Quality MonitoringSingle governance layer for data and AI assets across the platform
Data ManagementDelta Lake, Iceberg, bronze/silver/gold medallionOpen-format lakehouse storage with ACID transactions
CollaborationDelta Sharing, Marketplace, Clean RoomsSecure cross-organisation data sharing and collaboration

Serve & Analyse​

ComponentTypeConsumer
DashboardsAI/BIBusiness users
GenieAI/BI conversational analyticsBusiness users
Cloud BigTableOperational DB (low-latency NoSQL)Applications
Cloud SQLOperational DB (relational)Applications
Data StoreOperational DB (Firestore)Applications
LookerBI / semantic layerAnalysts, executives
Databricks AppsNative application hostingInternal apps
Business / AI AppCustom applicationEnd users
Data ConsumerSharing PartnerExternal partners

Integrate (cross-cutting)​

ComponentRole
ID ProviderAuthentication and SSO for platform access
Enterprise Catalog (Governance)External data catalog federation; policy import/export
AnthropicThird-party frontier model access via AI Gateway
Vertex AIGCP-native model registry, AutoML, managed endpoints
External OrchestratorAirflow, Cloud Composer, or other external schedulers triggering Databricks jobs

Storage​

Google Cloud Storage is the single underlying object store. All Delta Lake tables, Iceberg tables, and raw files persist in GCS buckets, decoupling compute from storage and enabling independent scaling of each tier.

Key Design Decisions​

  • Open lakehouse over proprietary lock-in. Delta Lake and Iceberg are open table formats. Unity Catalog governs both, and data can be read by any Iceberg-compatible engine, reducing dependence on a single vendor.
  • Unity Catalog as the single governance plane. All data assets — tables, ML models, volumes, notebooks — are registered in Unity Catalog. Access control is attribute-based and enforced at query time, not at the ETL layer.
  • Medallion architecture for data quality. Bronze (raw), silver (cleaned/conformed), gold (business-ready aggregates) layers enforce quality gates and make data lineage traceable end-to-end.
  • AI Gateway as the frontier model abstraction. All GenAI calls (Anthropic, Vertex AI, self-hosted models) are routed through the AI Gateway, enabling rate limiting, cost attribution, audit logging, and model swapping without application changes.
  • Google Cloud Storage as the decoupled storage tier. Compute (Databricks clusters) and storage (GCS) scale independently. All data survives cluster termination, and multiple compute engines can access the same tables concurrently.
  • Sharing via Delta Sharing, not data copy. External partners and marketplaces receive live, governed access to data shares rather than point-in-time extracts, keeping a single source of truth.