ML-PRO Sample Questions & Answers
Scaling and tuning Spark ML with advanced MLflow usage ties with validating and testing across the model lifecycle for the top weight, alongside deployment strategies built on custom model serving and environment architectures.
Launch the full ML-PRO simulator →Showing 10 of 20 free samples.
- Question 1Intermediate
Model Development · Advanced Feature Store Concepts
True or False: When using Databricks Feature Store, point-in-time correctness is automatically guaranteed for batch inference jobs without any specific configuration required in the
create_training_setmethod.Show answer & explanation
Correct answer: B
False. To ensure point-in-time correctness and prevent data leakage, you must provide a timestamp lookup key in your primary key DataFrame when calling
fs.create_training_set(). The Feature Store uses this timestamp to join features that were valid at that specific point in time, preventing the model from being trained on data that would not have been available at the time of prediction. - Question 2IntermediateSelect 2
Model Development · Scaling and Tuning
A large-scale image classification model is being trained on Databricks. The team observes that the training process is bottlenecked by the single-node driver's ability to coordinate the workers. They decide to explore distributed hyperparameter tuning to find optimal learning rates. Which two of the following technologies are natively integrated with Databricks for distributed hyperparameter tuning and can effectively manage this workload? (Select TWO)
Show answer & explanation
Correct answers: A, B
- Question 3Advanced
MLOps · Drift Detection and Lakehouse Monitoring
An ML team has configured Lakehouse Monitoring on an inference table. They receive an alert that the Population Stability Index (PSI) for a critical categorical feature, 'customer_segment', has exceeded the defined threshold. However, the drift analysis for the model's prediction and label columns shows no significant change. What is the most likely interpretation of this situation?
Show answer & explanation
Correct answer: B
A high PSI for an input feature indicates that the distribution of that feature has changed significantly between the baseline and current data (feature drift). The fact that prediction and label distributions are stable suggests that this change has not yet affected the model's output or the underlying data relationships. This could be because the feature has low importance, or the model is robust to this specific change. It serves as an early warning that requires investigation.
- Question 4Advanced
Model Development · Advanced Feature Store Concepts
A financial institution is building a real-time transaction fraud detection system. Latency is critical, as predictions must be returned in under 50 milliseconds. The features for this model require complex, on-the-fly calculations based on the user's recent activity, which is not available in the batch feature store. Which Databricks solution is best suited for this requirement?
Show answer & explanation
Correct answer: B
Databricks Feature Serving is designed for use cases that require ultra-low latency and on-demand feature computation. It allows you to define functions that compute features at inference time using data provided in the request. This avoids the need to pre-compute and store all possible feature values, making it ideal for real-time applications with dynamic feature requirements.
- Question 5Advanced
MLOps · Model Lifecycle Management
Case Study:
A global logistics company, ShipFast, wants to build an MLOps platform on Databricks to manage hundreds of models that predict package delivery times. Their key requirements are strict separation of development, staging, and production environments, auditable model transitions, and automated testing before any model is promoted to production.
The current process is manual, with data scientists promoting models via the UI, leading to inconsistent testing and accidental deployments. The MLOps team has been tasked with designing a fully automated, code-driven CI/CD pipeline. The pipeline must enforce that any model version proposed for the 'Staging' stage must first pass a suite of integration tests, including performance evaluation on a holdout dataset and a bias check. If the tests pass, the model version should be automatically transitioned to 'Staging' with a comment linking to the CI job results.
The MLOps team decides to use Databricks Jobs and Model Registry webhooks. They create a multi-task job that checks out the code, runs the tests, and on success, transitions the model. They need to ensure this job is triggered securely and reliably whenever a data scientist registers a new model version.
What is the most secure and robust way to architect the trigger mechanism for this validation pipeline?
Show answer & explanation
Correct answer: B
This approach is the most secure, integrated, and aligned with the requirements. Using a Job webhook directly links the Model Registry event to a Databricks Job without external middleware. Triggering on
TRANSITION_REQUEST_CREATEDinstead ofMODEL_VERSION_CREATEDcorrectly implements the business logic: the tests run when a promotion is requested, acting as a gatekeeper. This allows data scientists to register many experimental versions without triggering a pipeline for each one, only for those they wish to promote. - Question 6Intermediate
Model Development · Scaling and Tuning
A team is training a large language model and needs to parallelize the training process across multiple GPUs on a Databricks cluster. They are comparing Ray and Spark for this task. Which statement accurately describes a key difference between Ray and Spark in the context of distributed ML training?
Show answer & explanation
Correct answer: B
Ray is designed as a universal framework for distributed applications, offering fine-grained control over tasks and actors (stateful workers). This makes it highly suitable for the iterative and often stateful nature of modern ML training algorithms. Spark, conversely, is built around the RDD/DataFrame abstraction, which is excellent for data-parallel transformations (ETL) but can be less efficient for the communication-intensive patterns of distributed training.
- Question 7Intermediate
MLOps · Validation Testing
An ML engineer is designing an integration test for a complete machine learning pipeline. The pipeline involves feature engineering, model training, registration, and batch inference. The goal of the integration test is to verify that all components work together correctly. Which of the following is the MOST effective strategy for this integration test?
Show answer & explanation
Correct answer: C
The primary goal of an integration test is to ensure that different parts of a system work together. Using a small, controlled dataset makes the test fast and deterministic. The key assertions should be about the successful execution of the pipeline and the integrity of the final output (e.g., correct schema, non-null predictions), rather than model performance, which is typically evaluated in a separate model validation step.
- Question 8Advanced
Model Deployment · Custom Model Serving
You are deploying a custom model using the MLflow Deployments SDK. The model requires a specific set of GPU drivers and a proprietary C++ library to be installed in the serving environment. How can you specify these complex, non-PyPI dependencies for the model serving endpoint?
Show answer & explanation
Correct answer: B
For dependencies that cannot be installed via pip or conda, such as system-level libraries or drivers, Databricks Model Serving supports the use of custom Docker images. You can build a Docker image with all the required dependencies, push it to a container registry, and then specify the image URI in the endpoint configuration. This provides maximum flexibility for defining the serving environment.
- Question 9Beginner
MLOps · Drift Detection and Lakehouse Monitoring
When defining a monitor using Databricks Lakehouse Monitoring, what is the primary purpose of specifying a
snapshotanalysis type?Show answer & explanation
Correct answer: C
The
snapshotanalysis type is used for tables where data is not time-dependent, and each row represents an independent record (e.g., a table of customer profiles). The monitor analyzes the entire table as a single distribution and compares it against a baseline or previous analyses. This contrasts withTimeSeriesorInferenceLogwhich are designed for data with a timestamp component. - Question 10Intermediate
Model Development · Scaling and Tuning
A machine learning engineer is using the Pandas Function API (
applyInPandas) to train multiple time series forecasting models in parallel, one for each product category. What is the primary benefit of this approach compared to training models sequentially in a loop on the driver node?Show answer & explanation
Correct answer: C
The Pandas Function API, specifically
groupBy().applyInPandas(), allows you to apply a function (like model training) to each group of data in parallel across the Spark cluster's worker nodes. This is a powerful pattern for 'embarrassingly parallel' problems, such as training independent models for many different stores, products, or users. It drastically reduces overall training time by utilizing the full cluster.
Ready for the real thing?
The full ML-PRO simulator has every exam-style question, timed mode, and instant scoring.