DEA-C01 logo
Focused certification exam prep
Start practice

DEA-C01 Exam Domains 2026: Complete Guide to All 4 Content Areas

TL;DR
  • The AWS Certified Data Engineer - Associate (DEA-C01) guide has four domains weighted 34%, 26%, 22%, and 18% of scored content.
  • The exam has 65 questions (50 scored, 15 unscored) in 130 minutes; you need a scaled score of 720 out of 1,000.
  • Scoring is compensatory, so a weak domain can be offset by strength elsewhere, but Domain 1 carries the most weight.
  • The v1.1 guide covers 17 tasks and 120 skills, including vector indexes, Apache Iceberg, and LLM integration in data processing.

How the DEA-C01 Exam Is Built

The AWS Certified Data Engineer - Associate (DEA-C01) exam measures whether you can ingest, transform, store, operate, and secure data pipelines on AWS. The current published guide is Exam Guide v1.1, dated December 12, 2025. It organizes the exam into four content domains, 17 task statements, and 120 numbered skills. Understanding that hierarchy is the fastest way to turn a vague "learn AWS data services" goal into a concrete plan.

Each exam delivers 65 questions in 130 minutes. Of those, 50 are scored and 15 are unscored; AWS does not tell you which are which, so every question deserves full effort. Questions are multiple-choice (one correct answer, three distractors) or multiple-response (two or more correct answers among five or more options). Unanswered questions count as incorrect and there is no penalty for guessing, so never leave an item blank.

DomainWeightTasks
Data Ingestion and Transformation34%1.1 to 1.4
Data Store Management26%2.1 to 2.4
Data Operations and Support22%3.1 to 3.4
Data Security and Governance18%4.1 to 4.5
Scaled score, not a percentage: The passing mark is a scaled 720 on a 100 to 1,000 scale. It is not 72% of questions correct, and a practice-test percentage cannot be converted into the official score. For more on how this works, see our breakdown of the DEA-C01 passing score.

Because scoring is compensatory across the whole exam, there is no separate pass requirement per domain. That does not mean you can ignore a domain, though. The guide states its content list is not exhaustive, so thin knowledge in any area leaves you exposed to unfamiliar scenarios. For the full preparation picture, pair this article with the DEA-C01 study guide.

Domain 1: Data Ingestion and Transformation (34%)

This is the largest domain, and it deserves the largest share of your study hours. It spans four tasks: ingestion (1.1), transformation and processing (1.2), pipeline orchestration (1.3), and programming concepts (1.4). Domain 1 alone accounts for roughly a third of scored content, which makes it the single biggest lever on your result.

Task 1.1: Perform data ingestion

What the 12 skills cover

Ingestion questions test whether you can pick the right pattern for the source and arrival style, then keep it reliable under load.

  • Streaming reads from Kinesis, Amazon MSK, DynamoDB Streams, AWS DMS, AWS Glue, and Redshift (Skill 1.1.1).
  • Batch reads from S3, Glue, EMR, DMS, Redshift, Lambda, and Amazon AppFlow (Skill 1.1.2).
  • Schedulers and triggers: EventBridge, Apache Airflow, time-based jobs and crawlers, S3 Event Notifications, and Lambda invoked through Kinesis (Skills 1.1.5 to 1.1.7).
  • Throttling and rate limits across DynamoDB, RDS, and Kinesis (Skill 1.1.9).
  • Fan-in and fan-out for streams, plus replayability of ingestion pipelines (Skills 1.1.10 and 1.1.11).
  • Stateful versus stateless transactions (Skill 1.1.12).

Expect scenario questions that describe a source, a latency requirement, and a failure condition, then ask which combination of services fits. Knowing why you would pick Kinesis Data Streams over a batch pull, or how a replayable design recovers from a downstream failure, matters more than memorizing service names.

Task 1.2: Transform and process data

Here the guide leans toward practical engineering: connecting sources over JDBC and ODBC, integrating multiple sources, converting formats such as .csv to Apache Parquet, and optimizing processing cost. You must also be able to debug transformation failures and performance problems, build data APIs for other systems, and reason about the volume, velocity, and variety of data. The newer Skill 1.2.10 adds integrating large language models into processing. This is about wiring LLMs into a data pipeline as a processing step, not training or running inference on models, which the guide places out of scope.

Task 1.3: Orchestrate data pipelines

Orchestration questions compare Lambda, EventBridge, Amazon MWAA, Step Functions, and Glue workflows. The task also expects you to design for performance, availability, scalability, resiliency, and fault tolerance, to maintain serverless workflows, and to configure alerts through SNS and SQS. A common trap is choosing a heavyweight orchestrator when a simple event-driven Step Functions workflow meets the requirement.

Task 1.4: Apply programming concepts

Despite the listed languages (Python, SQL, Scala, R, Java, Bash, PowerShell), the exam is not a test of language-specific syntax. It tests engineering practice: version control, testing, logging, monitoring, infrastructure as code with CloudFormation and CDK, packaging serverless pipelines with AWS SAM, CI/CD concepts, Lambda concurrency and performance, mounting volumes within Lambda, and foundational ideas such as distributed computing, graphs, and trees.

Domain 2: Data Store Management (26%)

Domain 2 asks you to choose, organize, catalog, and age data appropriately. It has four tasks: choosing a data store (2.1), understanding data cataloging systems (2.2), managing the data lifecycle (2.3), and designing data models and schema evolution (2.4).

Task 2.1: Choose a data store

Matching store to access pattern

The guide expects cost-and-performance-aware selection across Redshift, EMR, Lake Formation, RDS, DynamoDB, Kinesis Data Streams, and MSK.

  • Use cases that match specific engines, such as HNSW indexing in Aurora PostgreSQL and fast key/value access in Amazon MemoryDB (Skill 2.1.3).
  • Redshift federated queries, materialized views, and Spectrum for remote or cross-store querying (Skill 2.1.5).
  • Locking behavior in Redshift and RDS (Skill 2.1.6).
  • Apache Iceberg as an open table format (Skill 2.1.7).
  • Vector index types, specifically HNSW and IVF (Skill 2.1.8).
  • AWS Transfer Family for migration and transfer integration (Skill 2.1.4).

Vector indexes and Iceberg are the newest additions and are frequent targets of candidate curiosity. Treat them as data-engineering topics: when is an index type appropriate, and how do open table formats change how you manage tables on S3? You do not need to train or serve models.

Task 2.2: Data cataloging systems

The central skills are building and referencing technical catalogs with the AWS Glue Data Catalog and Hive metastore, discovering schema and populating catalogs with Glue crawlers, synchronizing partitions, and connecting catalogs to sources and targets. Skill 2.2.6 introduces business catalogs through SageMaker Catalog, which differ in purpose from a technical catalog.

Task 2.3: Manage the lifecycle of data

Expect questions on loading and unloading between S3 and Redshift, S3 Lifecycle transitions and expiration, S3 versioning, DynamoDB TTL, deletion for business or legal requirements, and protecting data for resilience and availability. These are often cost-optimization scenarios in disguise.

Task 2.4: Data models and schema evolution

This task covers schema design in Redshift, DynamoDB, and Lake Formation, adapting to changing data characteristics, schema conversion, lineage tracking with SageMaker ML Lineage Tracking and SageMaker Catalog, and best practices for indexing, partitioning, and compression. Skill 2.4.6 adds vectorization concepts, referencing an Amazon Bedrock knowledge base.

Partitioning and compression are cross-cutting: Skill 2.4.5 shows up indirectly in Athena cost questions, Glue job tuning, and Redshift performance scenarios. Learn it once deeply and you will recognize it in many disguises.

Domain 3: Data Operations and Support (22%)

Domain 3 is about running pipelines day to day. Its four tasks are automating data processing (3.1), analyzing data (3.2), maintaining and monitoring pipelines (3.3), and ensuring data quality (3.4).

Automation and analysis (Tasks 3.1 and 3.2)

You will need to orchestrate with MWAA and Step Functions, troubleshoot managed workflows, use SDKs to access AWS features programmatically, prepare transformations with Glue DataBrew and SageMaker Unified Studio, and query with Athena. On the analysis side, the guide names visualization with DataBrew and QuickSight, data verification and cleaning, SQL queries and views in Redshift and Athena, exploration through Athena Spark notebooks, and the provisioned-versus-serverless tradeoff. Skill 3.2.6 covers aggregation, rolling averages, grouping, and pivoting as concepts.

Monitoring and troubleshooting (Task 3.3)

  • Extracting and analyzing logs with Athena, EMR, OpenSearch Service, and CloudWatch Logs Insights.
  • Configuring monitoring alerts and automating application logging with CloudWatch Logs.
  • Tracking API calls with CloudTrail.
  • Troubleshooting and maintaining Glue and EMR pipelines.

Data quality (Task 3.4)

This task checks that you can detect problems such as empty fields during processing, build quality rules and investigate consistency with DataBrew, describe sampling techniques, and handle data skew. Skew is a classic performance topic that appears in EMR and Glue tuning scenarios.

Key Takeaway

Domain 3 rewards operational thinking. When a question describes a failing or slow pipeline, identify whether the root cause is configuration, data shape (skew, volume), permissions, or monitoring gaps before choosing a service.

Domain 4: Data Security and Governance (18%)

The smallest domain still matters, and its concepts leak into the other three. It contains five tasks: authentication (4.1), authorization (4.2), encryption and masking (4.3), audit logs (4.4), and privacy and governance (4.5).

What to master

Security questions reward least-privilege reasoning and knowing which managed control fits which requirement.

  • Authentication: VPC security groups, IAM roles and endpoints, Secrets Manager credential creation and rotation, S3 Access Points, PrivateLink, and SageMaker Unified Studio domains, domain units, and projects.
  • Authorization: custom IAM policies when managed policies fall short, Parameter Store versus Secrets Manager, Redshift database authorization, Lake Formation permissions, and role-, tag-, and attribute-based approaches.
  • Encryption and masking: KMS keys, cross-account encryption, encrypting in and before transit, and anonymization to meet legal or policy needs.
  • Audit: CloudTrail, CloudTrail Lake, CloudWatch Logs, and analysis through Athena and OpenSearch Service, including large-volume EMR logs.
  • Governance: Redshift data sharing, PII identification with Macie and Lake Formation, preventing replication into prohibited Regions, AWS Config, and data sovereignty.

Because the exam is compensatory, strong security knowledge can offset a weaker area, but security is also easy points: the patterns repeat, and managed-service answers are usually preferred over custom code.

Quirks in the v1.1 Guide You Should Know About

Two inconsistencies in the published materials are worth understanding rather than ignoring.

  • AWS Schema Conversion Tool: The revision record for v1.1 removes AWS SCT from the in-scope service list, yet Skill 2.4.3 still names AWS SCT alongside AWS DMS Schema Conversion. Both observations are real. Prioritize DMS Schema Conversion, understand schema conversion concepts generally, and watch for any clarification from AWS.
  • Quick versus QuickSight: The service inventory says Amazon Quick, while Skills 3.2.1 and 3.2.2 say QuickSight. The guide uses both labels, so be ready to recognize either.

Also note that Amazon S3 Tables appears twice in the storage inventory. That is a duplicate listing, not extra scope. The guide states its content and service lists are non-exhaustive and may change, so always check the official exam guide before your test date.

What Is Out of Scope

AWS explicitly excludes machine-learning model training and inference, programming-language-specific syntax, and drawing business conclusions from data. Several named services are also out of scope, including AWS Amplify, AWS AppSync, AWS X-Ray, and AWS Elastic Beanstalk, among others on the official out-of-scope list. Do not spend time there.

The recommended background is two to three years in data engineering or data architecture and one to two years using AWS. This is guidance, not an admission rule; see our DEA-C01 requirements article for eligibility details. If you want a sense of how demanding the exam feels in practice, read how hard the DEA-C01 exam is.

Sequencing Your Preparation by Domain Weight

This is the one place for a schedule, and it is tied to the weights rather than generic advice. Start with Domain 1 because it is the heaviest and because ingestion, transformation, and orchestration concepts underpin the rest.

Weeks 1-3

Domain 1 first

  • Kinesis, MSK, DMS, Glue, and Lambda ingestion patterns and replayability.
  • Parquet conversion, Glue and EMR transformations, Step Functions versus MWAA.
  • IaC with CloudFormation, CDK, and SAM.
Weeks 4-5

Domain 2

  • Redshift, DynamoDB, Aurora, and MemoryDB selection by access pattern.
  • Glue Data Catalog, crawlers, partition sync, S3 Lifecycle, Iceberg, and vector index types.
Week 6

Domain 3

  • Athena, DataBrew, CloudWatch Logs Insights, data skew, and quality rules.
Week 7

Domain 4 and review

  • Lake Formation permissions, KMS, Secrets Manager, Macie, CloudTrail Lake.
  • Timed mixed-domain practice, then revisit weak tasks.

Adjust the pace to your background. A candidate with years of Spark and Airflow can compress Weeks 1 to 3, while someone newer to AWS security may stretch Week 7. A printable summary is available in the DEA-C01 cheat sheet, and you can test yourself on all four domains with the DEA-C01 practice tests.

Fees, Booking, and Exam-Day Mechanics

The exam is delivered through Pearson VUE, either at a test center or through online proctoring. The fee is USD 150, with applicable taxes. No separate member and non-member price schedule is published for this credential. An active AWS Certification can give a 50% discount benefit on a next certification exam, but that is a specific benefit, not a guaranteed price for every candidate. See the DEA-C01 certification cost breakdown for the details.

  • Languages: English, Japanese, Korean, and Simplified Chinese.
  • Accommodations: an eligible ESL +30-minute accommodation must be requested before booking.
  • Breaks: there are no scheduled breaks. Unscheduled breaks at a test center consume exam time, and online candidates cannot leave camera view without approval.
  • Validity: three years. Renewal is by passing the latest version of this exam; AWS's recertification table also lists passing AWS Certified Generative AI Developer - Professional.

Scheduling windows and deadlines are covered in DEA-C01 exam dates. AWS does not publish an issuer pass rate for this exam in the sources reviewed, which is discussed in DEA-C01 pass rate.

Who Hires for This Skill Set

The domains map directly onto roles that build and operate data platforms: data engineers, analytics engineers, cloud data architects, and platform engineers supporting analytics teams. Employers in any industry running data lakes or warehouses on AWS look for people comfortable with Glue, Redshift, Kinesis, Lake Formation, and the surrounding security controls. The certification signals familiarity with the AWS data stack, though it does not by itself prove hands-on competence. No credential-specific salary premium has been verified, so treat earnings claims cautiously and read the salary analysis and the ROI analysis with that in mind.

Frequently Asked Questions

How many domains does the DEA-C01 exam have?

Four: Data Ingestion and Transformation (34%), Data Store Management (26%), Data Operations and Support (22%), and Data Security and Governance (18%). Together they contain 17 task statements and 120 skills in Exam Guide v1.1.

Do I have to pass each domain separately?

No. Scoring is compensatory across the whole exam, and you need a scaled score of at least 720 on a 100 to 1,000 scale. Domain weighting still means ingestion and transformation influences your result most.

Does the exam test machine-learning model training?

No. ML model training and inference are out of scope. LLM integration in data processing, vector indexes, and Bedrock knowledge-base concepts appear only as data-engineering skills.

Is a specific certification or degree required before sitting for DEA-C01?

No prior certification, degree, or mandatory course is required. Two to three years in data engineering and one to two years on AWS are recommended, not compulsory. The minimum age is 13, with parent or guardian consent for ages 13 to 17.

Should I study for AWS SCT given the guide inconsistency?

Focus first on AWS DMS Schema Conversion and general schema-conversion concepts. SCT was removed from the in-scope service list in v1.1 but is still named in Skill 2.4.3, so check the official guide for updates and avoid over-investing.

For a deeper look at how all four content areas connect, revisit the complete DEA-C01 exam domains guide, and use the practice question bank to check your readiness against each domain.

Ready to pass your DEA-C01 exam?

Put this into practice with free DEA-C01 questions across every exam domain.