← Certifications

CDEP · Professional

Certified Data Engineering Professional

Validating mastery of modern batch, streaming, and lakehouse data engineering

Questions

29 multiple choice

Duration

75 minutes

Pass mark

80%

Fee

$35 per attempt

About this certification

The Certified Data Engineering Professional (CDEP) credential is issued by the American Computer Society to recognize practitioners who can design, implement, and operate production-grade data platforms at scale. It targets engineers who already work with distributed systems and need to demonstrate depth across the full lifecycle of data: modeling, ingestion, transformation, storage, orchestration, and governance. The credential emphasizes real-world architectural tradeoffs rather than rote tool usage, reflecting the responsibilities of a senior data engineer or platform lead in a modern organization.

CDEP holders are expected to understand data modeling paradigms spanning dimensional modeling (star and snowflake schemas, slowly changing dimensions), Data Vault 2.0 (hubs, links, satellites), and normalized relational design, and to choose the right approach for analytical, operational, and hybrid workloads. The exam probes the ability to reason about grain, conformed dimensions, and schema evolution in warehouses such as Snowflake, BigQuery, and Redshift, as well as in lakehouse environments where schema-on-read and schema enforcement coexist.

A substantial portion of the credential addresses pipeline engineering: batch processing with Spark and SQL-based transformation frameworks like dbt, streaming processing with Kafka, Flink, and Spark Structured Streaming, and the distributed processing concepts that underpin both — partitioning, shuffle, watermarking, and exactly-once versus at-least-once semantics. Candidates must be comfortable reasoning about latency, throughput, and consistency tradeoffs when designing pipelines that serve both analytical and near-real-time use cases.

Storage and table format literacy is core to the exam: candidates must understand columnar file formats such as Parquet and ORC, and open table formats — Apache Iceberg, Delta Lake, and Apache Hudi — including how they provide ACID transactions, time travel, schema evolution, and efficient upserts on top of object storage. The credential also covers orchestration tools such as Airflow, Dagster, and Prefect, and the operational disciplines of idempotency, backfills, SLAs, and failure recovery that keep pipelines trustworthy at scale.

Data quality, observability, and lineage are treated as first-class engineering concerns, covering frameworks like Great Expectations and dbt tests, anomaly detection, and column- and table-level lineage tooling such as OpenLineage. The final domain addresses governance, privacy, and cost management, including data cataloging, access controls, regulatory considerations such as GDPR and CCPA, and the FinOps practices needed to control the cost of cloud compute and storage at scale.

Earning the CDEP signals to employers that a candidate can be trusted to architect and operate mission-critical data infrastructure, mentor junior engineers, and make sound tradeoffs between performance, cost, reliability, and compliance. It is well suited to data engineers, platform engineers, and analytics engineers preparing for senior or staff-level roles.

Syllabus and exam weighting

Data Modeling and Warehousing

20%

Covers dimensional modeling, Data Vault 2.0, and normalized design, along with schema evolution and modeling tradeoffs in modern warehouses and lakehouses. Candidates must select appropriate modeling strategies for varied analytical and operational needs.

  • ▪Star and snowflake schema design
  • ▪Slowly changing dimensions (Types 1-3, 6)
  • ▪Data Vault 2.0: hubs, links, satellites
  • ▪Third normal form and normalized OLTP modeling
  • ▪Conformed dimensions and bus matrices
  • ▪Schema evolution and schema-on-read vs schema-on-write
  • ▪Grain determination and fact table design
  • ▪Modeling for semi-structured and nested data

Batch and Streaming Pipelines

18%

Focuses on building reliable data movement and transformation pipelines using both batch and streaming paradigms. Candidates must reason about latency, ordering, and delivery semantics.

  • ▪Batch ETL/ELT design patterns
  • ▪dbt-based transformation workflows
  • ▪Kafka producers, consumers, and topics
  • ▪Spark Structured Streaming and Flink windowing
  • ▪Watermarking and late-arriving data
  • ▪At-least-once vs exactly-once semantics
  • ▪Change data capture (CDC) pipelines
  • ▪Micro-batch vs continuous processing

Distributed Processing

15%

Examines the internals of distributed compute engines and how partitioning, shuffling, and resource management affect pipeline performance. Candidates must diagnose and optimize distributed job performance.

  • ▪Apache Spark architecture: driver, executors, stages
  • ▪Partitioning strategies and data skew
  • ▪Shuffle operations and join strategies (broadcast, sort-merge)
  • ▪Resource management (YARN, Kubernetes)
  • ▪Caching, persistence, and lazy evaluation
  • ▪Query optimization and physical plans
  • ▪Scaling considerations for large joins and aggregations

Storage and Table Formats

15%

Covers columnar file formats and open table formats that enable ACID transactions and efficient analytics on object storage. Candidates must choose formats appropriate to update frequency, latency, and interoperability needs.

  • ▪Parquet and ORC columnar storage internals
  • ▪Apache Iceberg: snapshots, manifests, partition evolution
  • ▪Delta Lake: transaction log and time travel
  • ▪Apache Hudi: copy-on-write vs merge-on-read
  • ▪Compaction and small-file management
  • ▪Partitioning and file layout for query performance
  • ▪Interoperability across lakehouse engines

Orchestration and Reliability

17%

Addresses the operational discipline required to run pipelines dependably in production, including scheduling, dependency management, and recovery from failure. Candidates must design pipelines that are idempotent and resilient to reprocessing.

  • ▪DAG design in Airflow, Dagster, or Prefect
  • ▪Idempotent pipeline design
  • ▪Backfill strategies and reprocessing windows
  • ▪SLA definition and alerting
  • ▪Retry policies and failure isolation
  • ▪Dependency sensors and cross-pipeline triggers
  • ▪Deployment and versioning of pipeline code

Data Quality, Governance, and Cost Management

15%

Covers the observability, lineage, governance, privacy, and cost-control practices needed to operate trustworthy and compliant data platforms. Candidates must apply controls that balance compliance, reliability, and cloud spend.

  • ▪Data quality testing with Great Expectations and dbt tests
  • ▪Anomaly detection and data observability platforms
  • ▪Column- and table-level lineage (e.g., OpenLineage)
  • ▪Data cataloging and metadata management
  • ▪Access control, masking, and encryption
  • ▪Regulatory compliance (GDPR, CCPA)
  • ▪FinOps and cloud cost optimization for compute and storage
  • ▪Data retention and archival policies

Learning outcomes

  • ✓Design dimensional, Data Vault, and normalized data models appropriate to a given analytical or operational workload
  • ✓Architect batch and streaming pipelines that meet defined latency, throughput, and consistency requirements
  • ✓Select and justify storage and table formats (Parquet, Iceberg, Delta Lake) for a given access pattern
  • ✓Build orchestrated, idempotent pipelines with robust backfill and failure-recovery strategies
  • ✓Implement data quality checks, observability, and lineage tracking across a data platform
  • ✓Apply governance, privacy, and cost-management controls to a production data platform

Exam format

Delivery
Online proctored exam, available on demand via remote proctoring software with ID verification.
Retakes
A $35 fee applies per attempt. Candidates must wait 14 days between attempts and are limited to a maximum of 3 attempts within any 12-month period.
Pass mark
80% of 29 scored questions. Results are graded instantly in MyACS with a domain-by-domain breakdown.

Maintaining the credential

  • ▪Certification is valid for 3 years from the date of issue
  • ▪Holders must accrue 60 continuing professional development (CPD) hours within the 3-year cycle
  • ▪CPD hours may be earned through approved training, conference sessions, publications, or teaching
  • ▪Recertification requires signing an updated ethics and professional conduct attestation
  • ▪Lapsed certifications beyond a 6-month grace period require a full re-examination

Recommended reading

Fundamentals of Data Engineering

O'Reilly Media

Comprehensive overview of the data engineering lifecycle, from ingestion through serving, by Joe Reis and Matt Housley.

The Data Warehouse Toolkit

Wiley

The definitive reference on dimensional modeling by Ralph Kimball and Margy Ross.

Designing Data-Intensive Applications

O'Reilly Media

Martin Kleppmann's foundational text on distributed systems, storage, and stream processing.

Building the Data Lakehouse

Technics Publications

Bill Inmon's treatment of lakehouse architecture and table format design principles.