Salesforce Decode
Salesforcedecode
Back to questions
IntegrationAdvanceddata-lakeiceberganalytics

Design Salesforce to data lake ingestion with Apache Iceberg

Real World Scenario

Data platform team wants open table format Iceberg fed from Salesforce for ML feature store.

Expected Answer

• CDC to Kafka to Spark streaming write Iceberg tables • Partition by ingestion date and OrgId • Schema evolution Iceberg handling Salesforce field adds • Soft delete column IsDeleted in lake models • Feature store reads curated Iceberg not raw • Lineage metadata Apache Atlas or Data Cloud • SLA freshness for ML training pipelines

Follow-Up Questions & Answers

Click to expand — each follow-up includes a direct, interview-ready answer

Main difference: use case and scale. CDC to Kafka to Spark streaming write Iceberg tables. Partition by ingestion date and OrgId. Pick based on your integration pattern and team capability. Lake ingestion decouples analytics from CRM API limits—architects own contract between lake and CRM freshness. Optimize for scale and operational observability.

Architect Perspective

Lake ingestion decouples analytics from CRM API limits—architects own contract between lake and CRM freshness.