D-DS-FN-23 Sample Questions & Answers
Advanced analytics theory, application and interpretation make up nearly half the weighting, alongside big-data tools and technology, initial data analysis, operationalizing and visualizing results, the analytics lifecycle, and the data scientist's role.
Launch the full D-DS-FN-23 simulator →Showing 10 of 20 free samples.
- Question 1Beginner
Initial Analysis of the Data · Basic R commands for data exploration
A data scientist is performing an initial analysis of a dataset in R. They want to quickly get a summary of the central tendency, dispersion, and distribution shape for a continuous numerical variable named
product_cost. Which R command would be most effective for this purpose?Show answer & explanation
Correct answer: C
The
summary()function in R is specifically designed to provide a quick statistical summary of a variable. For a numerical vector, it returns the minimum, 1st quartile, median, mean, 3rd quartile, and maximum values. This single command gives a concise overview of central tendency (mean, median), dispersion (quartiles, min/max), and distribution shape. - Question 2IntermediateSelect 2
Operationalizing an Analytics Project and Data Visualization Techniques · Building presentations for specific audiences
During a project presentation to senior executives, a data scientist needs to convey the potential return on investment (ROI) of a new predictive maintenance model. Which data visualization best practices should they employ? (Select TWO)
Show answer & explanation
Correct answers: B, D
- Question 3Intermediate
Advanced Analytics - Theory, Application, and Interpretation of Results for Eight Methods · Logistic Regression
A hospital wants to predict patient readmission risk. A data scientist builds a logistic regression model and a decision tree model. To compare their performance, they generate ROC curves for both. The Area Under the Curve (AUC) for the logistic regression is 0.85, and for the decision tree, it is 0.78. What does this comparison indicate?
Show answer & explanation
Correct answer: C
The AUC represents a model's ability to discriminate between positive and negative classes across all possible thresholds. A higher AUC (closer to 1.0) indicates better performance. An AUC of 0.85 means the logistic regression model is superior to the decision tree (AUC 0.78) in its ability to correctly rank a randomly chosen positive instance higher than a randomly chosen negative one. It does not directly state overall accuracy, which depends on a single chosen threshold.
- Question 4Intermediate
Advanced Analytics - Theory, Application, and Interpretation of Results for Eight Methods · K-means clustering
A data scientist is using K-means clustering to segment customers based on their purchasing behavior. They have run the algorithm with K=3 and K=5. Which method should be used to determine the optimal number of clusters (K) for the dataset?
Show answer & explanation
Correct answer: C
The Elbow method is a common heuristic for finding the optimal number of clusters in K-means. It involves running the algorithm for a range of K values and plotting the WSS for each. The point where the rate of decrease in WSS sharply slows, forming an 'elbow' in the plot, is considered a good estimate for the optimal K. Gini impurity is for decision trees, and Lift charts are for classification models.
- Question 5Advanced
Advanced Analytics for Big Data - Technology and Tools · Hadoop ecosystem and related product use cases
A financial institution is processing a massive stream of real-time transaction data. They need a tool within the Hadoop ecosystem that is specifically designed for distributed, real-time computation on large data streams. Which tool best fits this requirement?
Show answer & explanation
Correct answer: B
Apache Spark, and particularly its Spark Streaming library, is designed for scalable, high-throughput, fault-tolerant processing of live data streams. It processes data in micro-batches, providing near real-time capabilities. While tools like Pig and Hive are excellent for batch processing, and HBase is a NoSQL database, Spark is the premier choice in the Hadoop ecosystem for real-time stream computation.
- Question 6Intermediate
Advanced Analytics for Big Data - Technology and Tools · MapReduce and Apache Hadoop
What is the primary role of the NameNode in a Hadoop Distributed File System (HDFS) architecture?
Show answer & explanation
Correct answer: C
The NameNode is the centerpiece of the HDFS architecture. It acts as the master server and maintains the filesystem metadata, including the directory tree and the mapping of file blocks to the DataNodes where they are stored. It does not store the data itself but directs clients to the correct DataNodes. DataNodes store the data blocks, and TaskTrackers (in YARN, the NodeManager/ApplicationMaster) execute tasks.
- Question 7Intermediate
Advanced Analytics - Theory, Application, and Interpretation of Results for Eight Methods · Linear regression
A data scientist has built a linear regression model to predict house prices. The model's R-squared value is 0.82. Which statement accurately describes the meaning of this value?
Show answer & explanation
Correct answer: B
R-squared, or the coefficient of determination, is a statistical measure that represents the proportion of the variance for a dependent variable that's explained by an independent variable or variables in a regression model. An R-squared of 0.82 means that 82% of the fluctuations in house prices can be accounted for by the model's inputs.
- Question 8Intermediate
Advanced Analytics for Big Data - Technology and Tools · Advanced SQL methods
A data analyst needs to calculate the rolling 7-day average sales for each product in a large sales database using SQL. Which advanced SQL function is specifically designed for this type of calculation?
Show answer & explanation
Correct answer: C
Window functions perform calculations across a set of table rows that are somehow related to the current row. To calculate a rolling 7-day average, you would use
AVG(sales) OVER (PARTITION BY product_id ORDER BY sales_date ROWS BETWEEN 6 PRECEDING AND CURRENT ROW). This allows you to compute the average over a moving 'window' of rows without collapsing them like aGROUP BYclause would. - Question 9Advanced
Advanced Analytics - Theory, Application, and Interpretation of Results for Eight Methods · Naïve Bayesian classifiers
During the Model Planning phase of the Data Analytics Lifecycle, a team is deciding between using a Naïve Bayesian classifier and a Decision Tree for a fraud detection problem. The dataset has many features, and some of them are likely to be correlated. How would this correlation impact the choice of model?
Show answer & explanation
Correct answer: A
The 'Naïve' in Naïve Bayes comes from its core assumption that all features are conditionally independent given the class. If features are highly correlated, this assumption is violated, which can lead to poor probability estimates and reduced model performance. Decision Trees, on the other hand, do not make this assumption and can handle correlated features effectively by selecting the most informative feature at each split.
- Question 10Intermediate
Advanced Analytics - Theory, Application, and Interpretation of Results for Eight Methods · Time Series Analysis
You are analyzing monthly sales data for the past five years and notice a distinct pattern that repeats every 12 months. You need to build a model to forecast sales for the next year. Which analytical method is most appropriate for this task?
Show answer & explanation
Correct answer: C
Time Series Analysis is specifically designed for analyzing and forecasting data points collected over time. The presence of a repeating 12-month pattern indicates seasonality, a key component that time series models like SARIMA (Seasonal ARIMA) or Holt-Winters exponential smoothing are built to handle. Other methods like linear regression or clustering do not account for the temporal dependencies and seasonal patterns inherent in this data.
Ready for the real thing?
The full D-DS-FN-23 simulator has every exam-style question, timed mode, and instant scoring.