← Databricks Certified Generative AI Engineer Associate
DATABRICKS · objective · 17% of the exam
Data Preparation — Databricks Certified Generative AI Engineer Associate
The official DATABRICKS documentation our Data Preparation practice questions are cited to. Review the primary sources, then practise.
Official references for this objective
-
Databricks — Clean, Correct Change Data Capture with Spark Declarative Pipelines | Databricks
Spark Declarative Pipelines (SDP) addresses this with Auto CDC, a declarative API that abstracts away merge orchestration and correctness guarantees. Instead of implementing imperative patterns,
-
Databricks — Building Data Pipelines with Lakeflow Spark Declarative Pipelines | Databricks
Delta tables using the read_files SQL function (CSV, JSON, TXT, Parquet)
-
Databricks — From Day-Old Data to Real-Time Retail: Modernizing CDC With Spark Declarative Pipelines and Lakeflow | Databricks
legacy change data capture approaches that behave like batch pipelines, creating delays, data inconsistencies, and operational risk
-
Databricks — Sponsored by: Precisely | Architecting Agentic-Ready Data Pipelines with Databricks and Precisely | Databricks
enrichment—embedded directly into Databricks pipelines via native UDFs, unifying structured and unstructured data to ensure accuracy and reliability
-
including patterns for CDC, schema evolution, and hybrid batch/streaming workloads that can be applied immediately in production environments
-
Databricks — Databricks Apps: A Magic Wand for Driving Data Quality Adoption and Observability | Databricks
Combined with table-triggered workflows, checks run automatically as data changes without added operational complexity