DY0-001 Sample Questions & Answers
Model design, evaluation and exploratory analysis tie with supervised, unsupervised and deep learning for the top weight, alongside data infrastructure and MLOps, statistical foundations, and specialized NLP or computer-vision applications.
Launch the full DY0-001 simulator →Showing 10 of 20 free samples.
- Question 1Intermediate
Machine Learning · 3.2.2 Dimensionality Reduction
A data scientist is performing dimensionality reduction on a high-dimensional dataset for visualization purposes. The goal is to preserve the local structure and reveal underlying clusters in two dimensions. Which algorithm is most suitable for this task?
Show answer & explanation
Correct answer: B
t-SNE is a non-linear dimensionality reduction technique specifically designed for visualizing high-dimensional data in low-dimensional space (typically 2D or 3D). It excels at revealing the underlying structure of data, such as clusters, by preserving the local similarities between data points. PCA, in contrast, is a linear technique focused on maximizing variance and may not effectively separate clusters that are not linearly separable.
- Question 2Beginner
Mathematics and Statistics · 1.1.3 Data Distribution Properties
During an exploratory data analysis (EDA) of a dataset containing customer ages, a data scientist observes that the distribution is right-skewed. What does this indicate about the data?
Show answer & explanation
Correct answer: B
A right-skewed (or positively skewed) distribution has a long tail extending to the right. This is caused by a smaller number of high-value outliers pulling the mean to the right. In such a distribution, the typical relationship is Mean > Median > Mode. The median is less affected by outliers than the mean, so it remains closer to the bulk of the data.
- Question 3Intermediate
Operations and Processes · 4.3.2 Model Monitoring and Maintenance
A financial institution is using a gradient boosting model to detect fraudulent transactions. After deployment, the MLOps team notices a gradual decrease in the model's F1-score over several months. This phenomenon is commonly referred to as:
Show answer & explanation
Correct answer: C
Model drift, also known as concept drift, occurs when the statistical properties of the target variable, which the model is trying to predict, change over time in unforeseen ways. This causes the model, which was trained on historical data, to become less accurate as time passes. The gradual decrease in performance is a classic symptom of model drift, often caused by changes in fraudulent behavior patterns.
- Question 4Intermediate
Specialized Applications of Data Science · 5.1.1 Text Processing and Analysis
A data scientist needs to build a model to classify news articles into categories like 'Sports', 'Politics', and 'Technology'. The input data consists of the raw text of the articles. Which sequence of NLP techniques is most appropriate for preparing this text data for a machine learning model?
flowchart TD A[Start: Raw Text] --> B{Process} B --> C[Vectorization] C --> D[Model Training]Show answer & explanation
Correct answer: A
This represents a standard and effective pipeline for text classification. 1) Tokenization breaks the text into individual words (tokens). 2) Stop word removal eliminates common words ('the', 'a', 'is') that add little semantic value. 3) TF-IDF (Term Frequency-Inverse Document Frequency) vectorization converts the cleaned text into numerical vectors that represent the importance of each word in the context of the entire corpus, which is an ideal input for most machine learning classifiers.
- Question 5Beginner
Modeling, Analysis, and Outcomes · 2.2.2 Model Evaluation Metrics
In linear regression, the R-squared value represents the proportion of the variance in the dependent variable that is predictable from the independent variable(s).
Show answer & explanation
Correct answer: A
This is the correct definition of the R-squared (coefficient of determination) value. It is a statistical measure that provides a 'goodness of fit' for the model, indicating how much of the variability in the outcome data can be explained by the model's inputs.
- Question 6IntermediateSelect 2
Machine Learning · 3.1.1 Classification Algorithms
A hospital is analyzing patient readmission rates. The data science team wants to build a model to identify high-risk patients. The available data includes patient demographics, medical history, and lab results. The medical review board has mandated that the final model must be highly interpretable so that clinicians can understand the factors driving a specific prediction. Which TWO of the following models would be most appropriate choices? (Select TWO).
Show answer & explanation
Correct answers: A, C
Logistic Regression is a linear model that is highly interpretable. The coefficients of the model directly indicate the influence of each feature on the predicted outcome's log-odds, making it easy for clinicians to understand which factors are most significant.
Decision Trees are inherently interpretable, often called 'white-box' models. The tree structure provides a clear set of if-then rules that can be easily visualized and followed to understand how a specific prediction was made, which is ideal for clinical review.
- Question 7Intermediate
Mathematics and Statistics · 1.1.2 Probability and Distributions
The Central Limit Theorem states that, for a sufficiently large sample size, the sampling distribution of the sample mean will be approximately normally distributed, regardless of the shape of the population distribution.
Show answer & explanation
Correct answer: A
This statement is the correct definition of the Central Limit Theorem (CLT). It is a fundamental concept in statistics that allows for making inferences about a population mean based on a sample mean, even if the population itself is not normally distributed, provided the sample size is large enough (commonly n > 30).
- Question 8Advanced
Modeling, Analysis, and Outcomes · 2.1.2 Feature Engineering
A data scientist is working with a dataset that has several categorical features with high cardinality (e.g., 'ZIP Code' with thousands of unique values). Which encoding technique is generally MOST appropriate to use for these features in a tree-based model like a Random Forest?
Show answer & explanation
Correct answer: C
Target Encoding (or Mean Encoding) is highly effective for high-cardinality categorical features, especially with tree-based models. It replaces each category with the mean of the target variable for that category. This captures information about the target variable directly in the feature, avoiding the massive dimensionality increase of One-Hot Encoding and the arbitrary ordering problem of Label Encoding. Proper implementation requires careful handling (e.g., using cross-validation) to prevent data leakage.
- Question 9Intermediate
Specialized Applications of Data Science · 5.3.1 Time Series and Forecasting
A retail company is building a time-series model to forecast weekly sales for the next year. The historical sales data exhibits strong seasonality with peaks during holidays and a clear upward trend over the past five years. Which of the following models is best suited to capture both trend and seasonality?
Show answer & explanation
Correct answer: D
SARIMA is an extension of the ARIMA model specifically designed for time series data that contains a seasonal component. It includes additional parameters to model the seasonality alongside the trend and autoregressive components. Given the strong seasonality and trend in the sales data, SARIMA is the most appropriate and powerful choice among the options.
- Question 10Advanced
Machine Learning · 3.1.2 Regression Algorithms
Company Background
'AeroDynamics', a major airline, aims to optimize its pricing strategy for flights. They want to predict the passenger load factor (percentage of seats filled) for each flight up to 90 days in advance. This prediction will be a key input into their dynamic pricing engine.Current Situation
The airline currently uses a simple historical average model, which performs poorly due to changing market conditions, competitor pricing, and seasonal demand shifts. They have collected a rich dataset including historical booking data, flight schedules, aircraft types, origin-destination pairs, competitor prices scraped daily, and relevant economic indicators.Requirements & Constraints
- The model must predict a continuous value (load factor from 0.0 to 1.0).
- The model needs to be updated frequently to adapt to new booking patterns.
- Feature importance is crucial; the pricing strategy team needs to understand which factors most influence the load factor.
- The model should be robust to multicollinearity, as features like competitor prices and economic indicators are likely correlated.
- The solution should prevent overfitting, as the model will be trained on a vast amount of historical data.
Which machine learning approach is the MOST suitable for AeroDynamics' problem?
Show answer & explanation
Correct answer: C
Elastic Net is the best choice here because it addresses all key requirements. As a regression model, it predicts a continuous value. Its key advantage is the combination of L1 (Lasso) and L2 (Ridge) regularization. L2 regularization helps prevent overfitting and handles multicollinearity well. L1 regularization can perform automatic feature selection by shrinking some feature coefficients to zero, which directly addresses the need for understanding feature importance. This combination makes it a robust and interpretable choice for this complex business problem.
Ready for the real thing?
The full DY0-001 simulator has every exam-style question, timed mode, and instant scoring.