CAIA · Professional (specialist)
Certified AI Assurance Practitioner
Validating, governing, and monitoring AI systems across the model lifecycle
Questions
29 multiple choice
Duration
75 minutes
Pass mark
80%
Fee
$35 per attempt
About this certification
The Certified AI Assurance Practitioner (CAIA) credential is designed for professionals who sit at the intersection of machine learning engineering, risk management, and compliance. As organizations deploy models that influence lending decisions, hiring, medical triage, and content moderation, a growing class of specialists is needed who can independently verify that these systems behave as claimed, remain within acceptable risk bounds, and satisfy emerging legal obligations. CAIA certifies that a holder can perform this verification work with rigor.
The exam begins from technical fundamentals: how supervised, unsupervised, and generative models are trained, what evaluation metrics actually measure, and where benchmarks mislead. Candidates must understand precision-recall tradeoffs, calibration, out-of-distribution generalization, and the statistical basis for confidence intervals on reported metrics, because assurance work is only as credible as the measurement methodology behind it.
A substantial portion of the body of knowledge addresses fairness and robustness. Candidates are expected to apply group fairness metrics such as demographic parity and equalized odds, understand their mathematical incompatibilities, and design red-team exercises against large language models covering prompt injection, jailbreaks, data extraction, and adversarial perturbation. This reflects the reality that assurance is not a one-time certification event but an adversarial, ongoing discipline.
Documentation and lifecycle monitoring form the operational backbone of the credential. Candidates must be able to author and critique model cards, system cards, and datasheets for datasets, and to design MLOps monitoring pipelines that detect data drift, concept drift, and performance decay in production. Assurance without durable artifacts and continuous telemetry is considered incomplete under this body of knowledge.
Finally, CAIA anchors technical practice in governance. Candidates must be conversant with the NIST AI Risk Management Framework 1.0 and its Govern-Map-Measure-Manage functions, ISO/IEC 42001 requirements for an AI management system, the EU AI Act's risk-tiered obligations for prohibited, high-risk, limited-risk, and minimal-risk systems, and relevant US executive-branch guidance including OMB memoranda on federal agency AI use. This grounding allows holders to translate technical findings into audit-ready compliance evidence.
CAIA is intended as a mid-career specialist credential rather than an entry point. It assumes working familiarity with machine learning and statistics, and it rewards candidates who have hands-on exposure to model evaluation, bias auditing, or AI governance programs, whether from a data science, risk, legal, or internal-audit background.
Syllabus and exam weighting
Machine Learning Fundamentals for Assurance
15%Covers the technical foundations an assurance practitioner needs to critically evaluate ML systems, including model families, training pipelines, and the statistical basis of evaluation. Emphasis is on knowing where technical claims can be overstated or misapplied.
- ▪Supervised, unsupervised, and generative model families
- ▪Training/validation/test splits and cross-validation
- ▪Overfitting, underfitting, and regularization
- ▪Feature engineering and data leakage risks
- ▪Statistical significance and confidence intervals for metrics
- ▪Foundation models, fine-tuning, and transfer learning
- ▪Limitations of benchmark datasets
Model Evaluation and Benchmarking
18%Focuses on rigorous quantitative evaluation of model performance, including metric selection, calibration, and benchmark design. Practitioners learn to distinguish genuine capability improvements from benchmark leakage or cherry-picked results.
- ▪Classification metrics: precision, recall, F1, AUC-ROC, AUC-PR
- ▪Calibration and reliability diagrams
- ▪Regression and generative model evaluation metrics
- ▪Benchmark contamination and data leakage detection
- ▪Human evaluation protocols and inter-rater reliability
- ▪Out-of-distribution and stress-test evaluation design
- ▪Comparative benchmarking methodology and reporting standards
- ▪Statistical testing for model comparison
Fairness and Bias Measurement
17%Addresses quantitative fairness auditing across the ML lifecycle, from training data representativeness through outcome disparities. Candidates learn the mathematical tradeoffs among competing fairness definitions and how to select metrics appropriate to context.
- ▪Sources of bias: sampling, label, and measurement bias
- ▪Demographic parity, equalized odds, and predictive parity
- ▪Impossibility results among fairness metrics
- ▪Disparate impact analysis and the four-fifths rule
- ▪Bias mitigation: pre-processing, in-processing, post-processing techniques
- ▪Intersectional and subgroup fairness auditing
- ▪Fairness in generative and LLM outputs
Robustness and LLM Red-Teaming
17%Covers adversarial and stress-testing methods for classical models and large language models, including structured red-team exercises. Emphasis is placed on repeatable methodology rather than ad hoc probing.
- ▪Adversarial examples and perturbation-based attacks
- ▪Prompt injection and jailbreak techniques
- ▪Data extraction and membership inference attacks
- ▪Red-team planning, scoping, and reporting frameworks
- ▪Automated red-teaming and adversarial test generation
- ▪Robustness metrics under distribution shift
- ▪Content safety and harmful-output classification testing
Documentation and Transparency Artifacts
13%Focuses on the documentation practices that make assurance findings durable and auditable, including model cards, system cards, and dataset documentation. Candidates learn to evaluate documentation for completeness against recognized templates.
- ▪Model cards: structure and required elements
- ▪System cards for compound and agentic systems
- ▪Datasheets for datasets and data statements
- ▪Intended use, out-of-scope use, and known limitations disclosure
- ▪Version control and change logs for models and datasets
- ▪Documentation for third-party and vendor-supplied models
MLOps Monitoring, Drift, and Governance
20%Combines operational monitoring practices with formal governance frameworks. Candidates must connect technical drift-detection pipelines to the accountability structures required by NIST AI RMF, ISO/IEC 42001, and the EU AI Act.
- ▪Data drift and concept drift detection methods
- ▪Production performance monitoring and alerting thresholds
- ▪Incident response and model rollback procedures
- ▪NIST AI RMF 1.0: Govern, Map, Measure, Manage functions
- ▪ISO/IEC 42001 AI management system requirements
- ▪EU AI Act risk tiers and conformity assessment obligations
- ▪US executive-branch AI guidance and federal agency requirements
- ▪Third-party/vendor AI risk assessment
Learning outcomes
- ✓Design and execute model evaluation plans that go beyond aggregate accuracy to surface subgroup and distributional risk
- ✓Select and correctly interpret fairness metrics, including recognizing when metrics are mathematically incompatible
- ✓Plan and conduct structured red-teaming and adversarial robustness testing for LLM-based and classical ML systems
- ✓Author model cards, system cards, and data statements that meet documentation expectations of regulators and auditors
- ✓Build monitoring strategies to detect data drift, concept drift, and performance degradation in deployed models
- ✓Map an organization's AI portfolio to NIST AI RMF functions, ISO/IEC 42001 clauses, and EU AI Act risk tiers, and identify compliance gaps
Exam format
- Delivery
- Online proctored exam via remote webcam proctoring, available on demand; a secure testing center option is also offered in select regions.
- Retakes
- A $35 USD fee applies per attempt. Candidates who do not pass must wait 14 days before retaking the exam, and are limited to a maximum of 3 attempts within any rolling 12-month period.
- Pass mark
- 80% of 29 scored questions. Results are graded instantly in MyACS with a domain-by-domain breakdown.
Maintaining the credential
- ▪Certification is valid for 3 years from the date of passing the exam
- ▪Holders must earn and log 60 Continuing Professional Development (CPD) hours over the 3-year cycle, drawn from approved training, conference, or applied-project activities
- ▪A signed ethics and professional conduct attestation must be renewed at each recertification cycle
- ▪Recertification may be completed via CPD hours or by retaking and passing the current version of the exam
- ▪Lapsed certifications beyond a 6-month grace period require a full exam retake to reinstate
Recommended reading
NIST AI Risk Management Framework (AI RMF 1.0)
National Institute of Standards and Technology (NIST)
The foundational US framework organizing AI risk management into Govern, Map, Measure, and Manage functions; core reading for the governance domain.
ISO/IEC 42001:2023 — Artificial intelligence management system
International Organization for Standardization
The first international standard specifying requirements for an organizational AI management system; used directly in exam questions on governance structure.
Fairness and Machine Learning: Limitations and Opportunities
MIT Press
Barocas, Hardt, and Narayanan's widely used text covering fairness metrics, their tradeoffs, and legal context for algorithmic discrimination.
Interpretable Machine Learning
Leanpub (freely available online)
Christoph Molnar's practical reference on model interpretability techniques used throughout evaluation and documentation domains.