- What You Are Actually Sitting: AWS Certified Data Engineer - Associate
- Exam Mechanics: Format, Fee, Scoring, and Booking
- Domain 1: Data Ingestion and Transformation (34%)
- Domain 2: Data Store Management (26%)
- Domain 3: Data Operations and Support (22%)
- Domain 4: Data Security and Governance (18%)
- Scope Quirks Worth Knowing Before You Study
- Eligibility vs. Recommended Experience
- A Domain-Weighted Study Sequence
- Using Practice Questions Without Fooling Yourself
- After You Pass: Careers and Renewal
- Frequently Asked Questions
- The exam has 65 questions in 130 minutes; 50 are scored and 15 are unscored, with no way to tell which.
- Passing requires a scaled score of 720 on a 100-1,000 scale, not a 72% raw percentage.
- Data Ingestion and Transformation carries 34% of scored content, so start there and spend the most time there.
- Scoring is compensatory: a weak domain can be offset by strength elsewhere, but ignoring one is risky.
What You Are Actually Sitting: AWS Certified Data Engineer - Associate
The code DEA-C01 identifies AWS Certified Data Engineer - Associate, issued by AWS. If you are new to the terminology, the explainers on what DEA-C01 is and what DEA-C01 stands for cover the naming. This guide focuses on how to pass it, and it only draws on facts that apply to this AWS credential.
The current exam guide is DEA-C01 Exam Guide v1.1, published December 12, 2025, the latest revision in the published revision record. No separate precise go-live date was established, so check the official exam guide page before booking rather than assuming a transition date. The guide defines four content domains, 17 task statements, and 120 numbered skills. Those skills are your real syllabus. Everything below is organized around them.
Exam Mechanics: Format, Fee, Scoring, and Booking
| Item | DEA-C01 detail |
|---|---|
| Questions | 65 total: 50 scored, 15 unscored (not identified to you) |
| Time | 130 minutes |
| Question types | Multiple-choice (one correct answer, three distractors) and multiple-response (two or more correct answers among five or more options) |
| Passing score | Scaled 720 on a 100-1,000 scale |
| Fee | USD 150, plus applicable taxes |
| Delivery | Pearson VUE, at a test center or via online proctoring |
| Languages | English, Japanese, Korean, Simplified Chinese |
| Validity | Three years |
What the scoring model means for your strategy
Three scoring facts shape how you should take the exam:
- Unanswered items count as incorrect and there is no penalty for guessing. Answer every question, including ones you flag and cannot resolve.
- Scoring is compensatory across the whole exam. There is no separate pass requirement per domain, so a strong showing in Domain 1 can absorb a weaker Domain 4. That is not permission to skip a domain, because the unscored pool means you cannot know which questions count.
- 720 is a scaled score, not 72% raw. A practice-test percentage cannot be converted into your official result. The full explanation is in DEA-C01 Passing Score 2026.
Cost and booking
The fee is USD 150 plus applicable taxes, and no separate member or non-member price schedule is published for this credential. An active AWS Certification can give a 50% discount benefit toward a next exam, but that is a specific benefit with conditions, not a universal discount. The full accounting is in the DEA-C01 certification cost breakdown, and scheduling mechanics are in DEA-C01 Exam Dates 2026.
Domain 1: Data Ingestion and Transformation (34%)
This is the largest domain, so it deserves the largest share of your study hours. It has four tasks: ingestion, transformation and processing, orchestration, and programming concepts. For a broader view of all four areas, see DEA-C01 Exam Domains 2026.
Task 1.1: Perform data ingestion
Expect scenario questions that make you choose between streaming and batch and then configure the choice correctly.
- Streaming sources: Kinesis, Amazon MSK, DynamoDB Streams, DMS, Glue, and Redshift.
- Batch sources: S3, Glue, EMR, DMS, Redshift, Lambda, and AppFlow.
- Triggers and schedules: EventBridge, Apache Airflow, time-based jobs and crawlers, S3 Event Notifications, and invoking Lambda through Kinesis.
- Operational behavior: throttling and rate limits (DynamoDB, RDS, Kinesis), stream fan-in and fan-out, IP allowlists, and API-based consumption.
- Concepts: replayability of ingestion pipelines, and stateful versus stateless transactions.
Replayability is easy to overlook. Be ready to explain why a pipeline that can reprocess from a retained stream or from raw data in S3 recovers better from a bad transformation than one that cannot.
Task 1.2: Transform and process data
- Choose among EMR, Glue, Lambda, and Redshift for a given transformation requirement.
- Convert formats, such as .csv to Apache Parquet, and know why columnar formats help analytic query cost and speed.
- Connect sources over JDBC and ODBC, integrate multiple sources, and optimize processing cost.
- Debug transformation failures and performance problems.
- Optimize container performance on EKS and ECS, and build data APIs for other systems.
- Define data volume, velocity, and variety, and integrate LLMs into processing.
Task 1.3: Orchestrate data pipelines
- Know the orchestration options: Lambda, EventBridge, MWAA, Step Functions, and Glue workflows.
- Design for performance, availability, scalability, resiliency, and fault tolerance.
- Implement serverless workflows and configure alerts with SNS and SQS.
Task 1.4: Apply programming concepts
- Optimize runtime and Lambda concurrency and performance.
- Use infrastructure as code: CloudFormation, CDK, and SAM for packaging serverless pipelines (Lambda, Step Functions, DynamoDB tables).
- Explain CI/CD for data pipelines, version control, testing, logging, and monitoring.
- Mount and use volumes within Lambda, and explain distributed computing, graphs, and trees.
The skill list names Python, SQL, Scala, R, Java, Bash, and PowerShell, but the guide lists language-specific syntax as out of scope. You need to reason about engineering practices, not memorize syntax.
Domain 2: Data Store Management (26%)
The second-heaviest domain is about matching a workload to the right store and managing it over time.
Choosing a data store (Task 2.1)
- Select storage by cost and performance across Redshift, EMR, Lake Formation, RDS, DynamoDB, Kinesis Data Streams, and MSK, and configure it for the access pattern.
- Match specialized use cases: HNSW indexing in Aurora PostgreSQL for vector search, and MemoryDB for fast key/value access.
- Use Redshift federated queries, materialized views, and Spectrum to query remotely, and know the locking behavior in Redshift and RDS.
- Understand open table formats such as Apache Iceberg, vector index types (HNSW and IVF), and Transfer Family for migration.
Catalogs, lifecycle, and schema design (Tasks 2.2-2.4)
- Catalogs: Glue Data Catalog and Hive metastore as technical catalogs, Glue crawlers for schema discovery, partition synchronization, and SageMaker Catalog for business catalogs.
- Lifecycle: S3 Lifecycle for tiering and expiration, S3 versioning, DynamoDB TTL, Redshift load and unload with S3, and deletion for business or legal requirements.
- Modeling: schemas for Redshift, DynamoDB, and Lake Formation, adapting to changing data, lineage tracking, partitioning, indexing, compression, and vectorization concepts including Amazon Bedrock knowledge bases.
Domain 3: Data Operations and Support (22%)
This domain tests whether you can run pipelines after they are built.
Automate and analyze (Tasks 3.1-3.2)
- Orchestrate with MWAA and Step Functions, troubleshoot managed workflows, and automate through Lambda, EventBridge, and SDKs.
- Query with Athena, prepare data with Glue DataBrew and SageMaker Unified Studio, and explore with Athena Spark notebooks.
- Visualize and clean with DataBrew, QuickSight, Jupyter Notebooks, and SageMaker Data Wrangler.
- Compare provisioned and serverless tradeoffs, and define aggregation, rolling averages, grouping, and pivoting.
Monitor and ensure quality (Tasks 3.3-3.4)
- Log and monitor for audit traceability using CloudWatch Logs and CloudTrail, and analyze logs with Athena, EMR, OpenSearch Service, and CloudWatch Logs Insights.
- Troubleshoot Glue and EMR pipelines and configure monitoring alerts.
- Apply quality checks during processing, such as empty fields, create DataBrew quality rules, investigate consistency, and describe data-sampling techniques.
- Handle data skew, a recurring cause of slow distributed jobs.
Domain 4: Data Security and Governance (18%)
It is the smallest domain, but it is dense and tends to reward precise service knowledge.
- Authentication (4.1): VPC security groups, IAM roles and endpoints, Secrets Manager credential creation and rotation, S3 Access Points, PrivateLink, and SageMaker Unified Studio domains, domain units, and projects.
- Authorization (4.2): custom IAM policies when managed ones fall short, least privilege, role-, tag-, and attribute-based authorization, Lake Formation permissions across Redshift, EMR, Athena, and S3, and credential storage in Secrets Manager or Systems Manager Parameter Store.
- Encryption and masking (4.3): KMS, cross-account encryption, encryption in and before transit, and masking or anonymization to satisfy law or policy.
- Audit logs (4.4): CloudTrail, CloudTrail Lake for centralized queries, CloudWatch Logs, and large-volume EMR logs.
- Privacy and governance (4.5): Macie and Lake Formation for PII, preventing replication into prohibited Regions, AWS Config for account-configuration changes, data sovereignty, Redshift sharing permissions, and SageMaker Catalog project access.
Scope Quirks Worth Knowing Before You Study
A few details in the published materials are inconsistent or easy to misread. Handling them calmly saves study time.
- AWS SCT versus DMS Schema Conversion. The v1.1 revision record removes AWS Schema Conversion Tool from the in-scope service list, yet Skill 2.4.3 still names SCT next to DMS Schema Conversion. Both observations are real. Learn DMS Schema Conversion well, treat SCT as lower priority, and do not assume the question is resolved either way.
- Quick versus QuickSight. The service inventory says Amazon Quick, while skills 3.2.1 and 3.2.2 say QuickSight. Study the visualization capability, and do not rely on an undocumented equivalence between the two labels.
- Lists are non-exhaustive. The guide states its content and service lists are not exhaustive and can change, so check the official in-scope and out-of-scope service pages near your exam date.
- What is out of scope. Machine-learning model training and inference, language-specific syntax, and drawing business conclusions from data are excluded. LLM integration, vector indexes, and Bedrock knowledge-base concepts appear only as data-engineering skills, not as ML assessment.
Eligibility vs. Recommended Experience
There is no required prior certification, degree, mandatory course, or work-hour total. The guide describes a target candidate with two to three years in data engineering or data architecture and one to two years of hands-on AWS experience, but that is guidance, not an admission rule. The general minimum age is 13, with parent or guardian consent required for ages 13-17. Details are in DEA-C01 Requirements 2026.
The recommended background does tell you where to invest if you are short on experience: ETL pipelines, source control with Git, data lakes, core networking, storage and compute, SQL, and data-quality thinking. If you are wondering how that translates into effort, How Hard Is the DEA-C01 Exam? gives a candid view. No issuer pass rate has been published; see DEA-C01 Pass Rate 2026 for what can and cannot be said.
A Domain-Weighted Study Sequence
Generic schedules are less useful than ordering your weeks by exam weight and dependency. Domain 1 comes first because its services (Glue, Kinesis, Lambda, Step Functions) recur in every other domain. Adjust the length to your background.
Domain 1: ingestion, transformation, orchestration
- Build one batch pipeline (S3 to Glue to Parquet) and one streaming pipeline (Kinesis to Lambda).
- Compare Step Functions, MWAA, and Glue workflows on the same job.
- Practice throttling, replay, and fan-out scenarios.
Domain 2: stores, catalogs, lifecycle
- Work through Redshift, DynamoDB, and Aurora use cases, plus Iceberg and vector index concepts.
- Configure Glue crawlers, partition sync, and S3 Lifecycle rules.
Domain 3: operations and quality
- Query logs with Athena and CloudWatch Logs Insights.
- Build DataBrew quality rules and review data skew.
Domain 4: security and governance
- Write least-privilege policies and configure Lake Formation permissions.
- Walk through KMS cross-account cases and Macie PII discovery.
Review and timed practice
- Sit full 130-minute simulations and review wrong answers against the skill list.
- Use the DEA-C01 cheat sheet for final recall.
Key Takeaway
Hands-on labs beat passive reading for Domain 1 because the exam asks you to choose and configure services under constraints. Build each pipeline type once, then explain why you did not pick the alternatives.
For a fuller walkthrough of resources and pacing, the main DEA-C01 study guide and the hub on DEA-C01 training are good companions.
Using Practice Questions Without Fooling Yourself
Practice banks are useful for pattern recognition, but they have limits you should respect:
- Vendor claims about bank size, coverage, or pass assurance are not issuer facts. Verify that any bank is aligned to the current code and the v1.1 guide.
- A practice percentage does not map to the 720 scaled score, so use it as a trend, not a prediction.
- Review every missed question by tracing it to a task or skill ID in the guide, then fix the gap with a lab or documentation read.
- Multiple-response items are where candidates lose points. Practice identifying all required answers rather than the first plausible two.
You can work through scenario-style questions on the main practice test site and keep a log of which domains your misses cluster in. Treat any bank as knowledge preparation, not as the real exam and not as proof of practical ability.
After You Pass: Careers and Renewal
The credential targets roles centered on building and operating data pipelines on AWS, such as data engineer, analytics engineer, and cloud data architect. Browse the types of roles in DEA-C01 jobs. No verified credential-specific salary premium or issuer pass rate was found, so treat compensation claims with caution; the salary guide and the ROI analysis evaluate the evidence rather than promising outcomes. A comparison with the Solutions Architect Associate (SAA-C03) is useful if you are deciding which Associate credential suits your role, since SAA-C03 is a separate credential with its own objectives.
Renewal
The certification is valid for three years. You can renew by passing the latest version of the Data Engineer Associate exam, and the current AWS Recertification table also permits passing AWS Certified Generative AI Developer - Professional to renew an active Data Engineer Associate certification. The renewed period runs from the date you complete the recertification action, not from the old expiry date. This credential's recertification row lists no CEU/CPE quota and no Skill Builder maintenance route, so do not borrow options from other certifications.
Frequently Asked Questions
There are 65 questions in 130 minutes. Fifty are scored and 15 are unscored, and you cannot tell which are which, so treat every question as if it counts.
A minimum scaled score of 720 on a 100-1,000 scale. It is not a 72% raw-score threshold, and a practice-test percentage cannot be converted into the official scaled score.
Data Ingestion and Transformation, which carries 34% of scored content. Its services also appear throughout the other three domains, so it builds the foundation for the rest.
No. AWS eligibility policy sets no prerequisite certification, degree, course, or work-hour total. The two-to-three years of data engineering and one-to-two years of AWS experience is recommended guidance only.
The fee is USD 150 plus applicable taxes. You take it through Pearson VUE at a test center or with online proctoring. An active AWS Certification may provide a 50% discount toward a next exam under AWS terms.