MLA-C01 Sample Questions & Answers
Ingesting data and engineering features for machine learning carries the top share, just ahead of monitoring, maintaining and securing deployed solutions, choosing and training models, and orchestrating ML workflows through CI/CD pipelines.
Launch the full MLA-C01 simulator →Showing 8 of 17 free samples.
- Question 1Intermediate
Data Preparation for Machine Learning (ML) · 1.1: Ingest and store data
A financial services company is building a real-time fraud detection system. The system ingests transaction data via Amazon Kinesis Data Streams. A Machine Learning Engineer needs to perform feature engineering on this streaming data (such as calculating rolling averages of transaction amounts over the last 10 minutes) before passing the features to a SageMaker endpoint for inference. Which approach offers the low-latency feature calculation required?
Show answer & explanation
Correct answer: B
Amazon Managed Service for Apache Flink allows for stateful processing of streaming data, such as calculating rolling windows, with very low latency. Writing the results to the SageMaker Feature Store (Online Store) makes these features immediately available for low-latency inference lookup.
- Question 2Advanced
Data Preparation for Machine Learning (ML) · 1.1: Ingest and store data
A Machine Learning Engineer is preparing a large dataset (50 TB) stored in Amazon S3 for training a computer vision model. The training job will run on a cluster of Amazon EC2 p4d.24xlarge instances using Amazon SageMaker. The dataset consists of millions of small image files. The engineer observes that the training job initialization is taking a long time due to S3 API latency when listing and downloading objects. Which storage configuration will MAXIMIZE data loading performance?
Show answer & explanation
Correct answer: C
Amazon FSx for Lustre is a high-performance file system optimized for fast processing of workloads such as machine learning and high performance computing (HPC). When linked to an S3 bucket, it lazy-loads data and presents it as a file system, significantly reducing the overhead of S3 API calls (LIST/GET) for millions of small files, maximizing GPU utilization.
- Question 3Intermediate
Data Preparation for Machine Learning (ML) · 1.3: Ensure data integrity and prepare data for modeling
A retail company wants to train a customer churn prediction model. The dataset contains sensitive Personally Identifiable Information (PII) including customer names and email addresses. The company's security policy requires that all PII be redacted before the data is used for model training. Which solution provides the MOST automated and serverless way to identify and redact this PII in Amazon S3?
Show answer & explanation
Correct answer: B
AWS Glue DataBrew is a visual data preparation tool that includes built-in transformations to detect and handle PII. It can automatically detect sensitive data types (like emails and names) and apply masking or redaction transformations (e.g., replacement, hashing) without writing code.
- Question 4Beginner
Data Preparation for Machine Learning (ML) · 1.2: Transform data and perform feature engineering
An ML Engineer is setting up a feature engineering pipeline. The requirement is to have a centralized repository where features can be stored, discovered, and shared across different teams. The solution must support both low-latency retrieval for real-time inference and high-throughput retrieval for batch training. Which AWS service component should be used?
Show answer & explanation
Correct answer: B
Amazon SageMaker Feature Store is a purpose-built repository for ML features. It supports an Online Store (low latency for inference, backed by DynamoDB) and an Offline Store (high throughput for training, backed by S3), keeping them synchronized.
- Question 5Beginner
ML Model Development · 2.1: Choose a modeling approach
A media company wants to analyze user reviews using Natural Language Processing (NLP). They need to categorize reviews into 'Positive', 'Negative', and 'Neutral' sentiments. The team has limited ML expertise and wants to avoid training and managing a custom model. Which AWS service is the BEST choice?
Show answer & explanation
Correct answer: B
Amazon Comprehend is a fully managed NLP service that uses machine learning to uncover information in text. It includes pre-trained capabilities for sentiment analysis, requiring no ML expertise to use.
- Question 6Intermediate
ML Model Development · 2.2: Train and refine models
An ML Engineer is training an XGBoost model using Amazon SageMaker. The training dataset is highly imbalanced, with only 2% of the records belonging to the positive class. The engineer wants to ensure the model learns to identify the positive class effectively. Which hyperparameter or configuration should be adjusted?
Show answer & explanation
Correct answer: C
In XGBoost, the 'scale_pos_weight' parameter controls the balance of positive and negative weights. For imbalanced datasets, setting this value (typically as sum(negative instances) / sum(positive instances)) helps the algorithm pay more attention to the minority class.
- Question 7Intermediate
Deployment and Orchestration of ML Workflows · 3.3: Set up CI/CD pipelines
A company is using Amazon SageMaker Pipelines to orchestrate an end-to-end ML workflow. The pipeline includes a training step followed by a model evaluation step. The company wants to automatically register the model in the SageMaker Model Registry ONLY if the model's accuracy on the evaluation dataset exceeds 90%. Which SageMaker Pipelines step should be used to implement this logic?
Show answer & explanation
Correct answer: B
The ConditionStep in SageMaker Pipelines allows the definition of conditional logic. It can evaluate the output of the evaluation step (e.g., accuracy > 0.90) and execute the registration step only if the condition evaluates to true.
- Question 8Intermediate
Deployment and Orchestration of ML Workflows · 3.1: Select deployment infrastructure
A startup is deploying a Large Language Model (LLM) for a chatbot application. The traffic pattern is highly unpredictable: there are long periods of inactivity followed by sudden bursts of user requests. The model size is 10 GB. Cost optimization is a primary concern, and cold-start latency of a few seconds is acceptable for the first request in a burst. Which SageMaker Inference option is MOST cost-effective?
Show answer & explanation
Correct answer: C
SageMaker Serverless Inference is designed for workloads with intermittent or unpredictable traffic. It automatically provisions compute capacity based on the volume of inference requests and scales down to zero when idle, charging only for the compute time used. The tolerance for cold starts aligns with Serverless Inference characteristics.
Ready for the real thing?
The full MLA-C01 simulator has every exam-style question, timed mode, and instant scoring.