CDEP · Professional
Certified Data Engineering Professional
Validating mastery of modern batch, streaming, and lakehouse data engineering
Questions
29 multiple choice
Duration
75 minutes
Pass mark
80%
Fee
$35 per attempt
About this certification
The Certified Data Engineering Professional (CDEP) credential is issued by the American Computer Society to recognize practitioners who can design, implement, and operate production-grade data platforms at scale. It targets engineers who already work with distributed systems and need to demonstrate depth across the full lifecycle of data: modeling, ingestion, transformation, storage, orchestration, and governance. The credential emphasizes real-world architectural tradeoffs rather than rote tool usage, reflecting the responsibilities of a senior data engineer or platform lead in a modern organization.
CDEP holders are expected to understand data modeling paradigms spanning dimensional modeling (star and snowflake schemas, slowly changing dimensions), Data Vault 2.0 (hubs, links, satellites), and normalized relational design, and to choose the right approach for analytical, operational, and hybrid workloads. The exam probes the ability to reason about grain, conformed dimensions, and schema evolution in warehouses such as Snowflake, BigQuery, and Redshift, as well as in lakehouse environments where schema-on-read and schema enforcement coexist.
A substantial portion of the credential addresses pipeline engineering: batch processing with Spark and SQL-based transformation frameworks like dbt, streaming processing with Kafka, Flink, and Spark Structured Streaming, and the distributed processing concepts that underpin both — partitioning, shuffle, watermarking, and exactly-once versus at-least-once semantics. Candidates must be comfortable reasoning about latency, throughput, and consistency tradeoffs when designing pipelines that serve both analytical and near-real-time use cases.
Storage and table format literacy is core to the exam: candidates must understand columnar file formats such as Parquet and ORC, and open table formats — Apache Iceberg, Delta Lake, and Apache Hudi — including how they provide ACID transactions, time travel, schema evolution, and efficient upserts on top of object storage. The credential also covers orchestration tools such as Airflow, Dagster, and Prefect, and the operational disciplines of idempotency, backfills, SLAs, and failure recovery that keep pipelines trustworthy at scale.
Data quality, observability, and lineage are treated as first-class engineering concerns, covering frameworks like Great Expectations and dbt tests, anomaly detection, and column- and table-level lineage tooling such as OpenLineage. The final domain addresses governance, privacy, and cost management, including data cataloging, access controls, regulatory considerations such as GDPR and CCPA, and the FinOps practices needed to control the cost of cloud compute and storage at scale.
Earning the CDEP signals to employers that a candidate can be trusted to architect and operate mission-critical data infrastructure, mentor junior engineers, and make sound tradeoffs between performance, cost, reliability, and compliance. It is well suited to data engineers, platform engineers, and analytics engineers preparing for senior or staff-level roles.
Syllabus and exam weighting
Data Modeling and Warehousing
20%Covers dimensional modeling, Data Vault 2.0, and normalized design, along with schema evolution and modeling tradeoffs in modern warehouses and lakehouses. Candidates must select appropriate modeling strategies for varied analytical and operational needs.
- ▪Star and snowflake schema design
- ▪Slowly changing dimensions (Types 1-3, 6)
- ▪Data Vault 2.0: hubs, links, satellites
- ▪Third normal form and normalized OLTP modeling
- ▪Conformed dimensions and bus matrices
- ▪Schema evolution and schema-on-read vs schema-on-write
- ▪Grain determination and fact table design
- ▪Modeling for semi-structured and nested data
Batch and Streaming Pipelines
18%Focuses on building reliable data movement and transformation pipelines using both batch and streaming paradigms. Candidates must reason about latency, ordering, and delivery semantics.
- ▪Batch ETL/ELT design patterns
- ▪dbt-based transformation workflows
- ▪Kafka producers, consumers, and topics
- ▪Spark Structured Streaming and Flink windowing
- ▪Watermarking and late-arriving data
- ▪At-least-once vs exactly-once semantics
- ▪Change data capture (CDC) pipelines
- ▪Micro-batch vs continuous processing
Distributed Processing
15%Examines the internals of distributed compute engines and how partitioning, shuffling, and resource management affect pipeline performance. Candidates must diagnose and optimize distributed job performance.
- ▪Apache Spark architecture: driver, executors, stages
- ▪Partitioning strategies and data skew
- ▪Shuffle operations and join strategies (broadcast, sort-merge)
- ▪Resource management (YARN, Kubernetes)
- ▪Caching, persistence, and lazy evaluation
- ▪Query optimization and physical plans
- ▪Scaling considerations for large joins and aggregations
Storage and Table Formats
15%Covers columnar file formats and open table formats that enable ACID transactions and efficient analytics on object storage. Candidates must choose formats appropriate to update frequency, latency, and interoperability needs.
- ▪Parquet and ORC columnar storage internals
- ▪Apache Iceberg: snapshots, manifests, partition evolution
- ▪Delta Lake: transaction log and time travel
- ▪Apache Hudi: copy-on-write vs merge-on-read
- ▪Compaction and small-file management
- ▪Partitioning and file layout for query performance
- ▪Interoperability across lakehouse engines
Orchestration and Reliability
17%Addresses the operational discipline required to run pipelines dependably in production, including scheduling, dependency management, and recovery from failure. Candidates must design pipelines that are idempotent and resilient to reprocessing.
- ▪DAG design in Airflow, Dagster, or Prefect
- ▪Idempotent pipeline design
- ▪Backfill strategies and reprocessing windows
- ▪SLA definition and alerting
- ▪Retry policies and failure isolation
- ▪Dependency sensors and cross-pipeline triggers
- ▪Deployment and versioning of pipeline code
Data Quality, Governance, and Cost Management
15%Covers the observability, lineage, governance, privacy, and cost-control practices needed to operate trustworthy and compliant data platforms. Candidates must apply controls that balance compliance, reliability, and cloud spend.
- ▪Data quality testing with Great Expectations and dbt tests
- ▪Anomaly detection and data observability platforms
- ▪Column- and table-level lineage (e.g., OpenLineage)
- ▪Data cataloging and metadata management
- ▪Access control, masking, and encryption
- ▪Regulatory compliance (GDPR, CCPA)
- ▪FinOps and cloud cost optimization for compute and storage
- ▪Data retention and archival policies
Learning outcomes
- ✓Design dimensional, Data Vault, and normalized data models appropriate to a given analytical or operational workload
- ✓Architect batch and streaming pipelines that meet defined latency, throughput, and consistency requirements
- ✓Select and justify storage and table formats (Parquet, Iceberg, Delta Lake) for a given access pattern
- ✓Build orchestrated, idempotent pipelines with robust backfill and failure-recovery strategies
- ✓Implement data quality checks, observability, and lineage tracking across a data platform
- ✓Apply governance, privacy, and cost-management controls to a production data platform
Exam format
- Delivery
- Online proctored exam, available on demand via remote proctoring software with ID verification.
- Retakes
- A $35 fee applies per attempt. Candidates must wait 14 days between attempts and are limited to a maximum of 3 attempts within any 12-month period.
- Pass mark
- 80% of 29 scored questions. Results are graded instantly in MyACS with a domain-by-domain breakdown.
Maintaining the credential
- ▪Certification is valid for 3 years from the date of issue
- ▪Holders must accrue 60 continuing professional development (CPD) hours within the 3-year cycle
- ▪CPD hours may be earned through approved training, conference sessions, publications, or teaching
- ▪Recertification requires signing an updated ethics and professional conduct attestation
- ▪Lapsed certifications beyond a 6-month grace period require a full re-examination
Recommended reading
Fundamentals of Data Engineering
O'Reilly Media
Comprehensive overview of the data engineering lifecycle, from ingestion through serving, by Joe Reis and Matt Housley.
The Data Warehouse Toolkit
Wiley
The definitive reference on dimensional modeling by Ralph Kimball and Margy Ross.
Designing Data-Intensive Applications
O'Reilly Media
Martin Kleppmann's foundational text on distributed systems, storage, and stream processing.
Building the Data Lakehouse
Technics Publications
Bill Inmon's treatment of lakehouse architecture and table format design principles.