DP-100 Sample Questions

DP-100 Sample Questions & Answers

Training and deploying models ties with optimizing language models through prompt flow and RAG for the top weight, alongside designing a solution and managing workspace assets, and running experiments with automated ML and notebooks.

Launch the full DP-100 simulator →

Showing 10 of 20 free samples.

  1. Question 1Intermediate

    Explore data, and run experiments · Use automated machine learning to explore optimal models

    A hospital is using an AutoML for tabular data job to predict patient readmission risk. The Responsible AI dashboard for the best model reveals that the model has a significantly lower prediction accuracy for a minority demographic group compared to the majority group. This indicates a potential fairness issue. What is the most appropriate first step to mitigate this bias?

    Show answer & explanation

    Correct answer: C

    Performance disparities often arise from imbalanced data where the model has insufficient examples for minority groups. The most effective initial step is to address this root cause by collecting more representative data or using sampling techniques (like SMOTE or random oversampling) to create a more balanced training dataset. This allows the model to learn the patterns for the minority group more effectively.

  2. Question 2Intermediate

    Train and deploy models · Manage models

    An MLOps engineer is packaging a trained MLflow model for deployment. To ensure seamless inference, they must include information about the required input data schema directly within the model artifacts. Which file within the MLflow model directory should be modified to include this schema definition?

    Show answer & explanation

    Correct answer: D

    The MLmodel file is the primary metadata file for an MLflow model. It contains essential information, including the model's flavor, creation time, and most importantly, the signature. The signature explicitly defines the input and output schemas (data types, names, and tensor shapes), which is used by deployment tools to validate requests and format data correctly for inference.

  3. Question 3Beginner

    Design and prepare a machine learning solution · Create and manage compute targets

    True or False: When using an Azure Machine Learning compute cluster for training, you are only billed for the compute time when a job is actively running on the nodes.

    Show answer & explanation

    Correct answer: A

    This is true. Azure ML compute clusters can be configured with a minimum number of nodes (typically zero). The cluster automatically scales up when a job is submitted and scales down to the minimum count when idle. You are only charged for the nodes when they are allocated and running, not for the cluster resource itself when it's scaled to zero nodes.

  4. Question 4Intermediate

    Train and deploy models · Deploy a model

    A team has deployed a machine learning model to a managed online endpoint with a blue-green deployment strategy. 80% of the traffic is currently directed to the stable 'blue' deployment, and 20% is directed to the new 'green' deployment for testing. After monitoring, the team confirms the green deployment is performing well and decides to route all traffic to it. Which command should be used to achieve this without causing downtime?

    Show answer & explanation

    Correct answer: B

    The az ml online-endpoint update command is used to modify the properties of an existing endpoint, including the traffic allocation. To route all traffic to the green deployment, you would use this command with the --traffic "green=100" parameter. This updates the endpoint's routing rules in place, ensuring a smooth transition with no downtime.

  5. Question 5Intermediate

    Explore data, and run experiments · Use notebooks for custom model training

    A data scientist is working on a local machine with the Azure Machine Learning SDK and needs to access data stored in an Azure Blob Storage container for interactive analysis in a Jupyter notebook. The workspace is configured with a datastore named blob_datastore. Which code snippet correctly accesses a file named data/customers.csv from this datastore as a Pandas DataFrame?

    Show answer & explanation

    Correct answer: A

    The azureml:// URI scheme is the standard way to reference data assets and paths within datastores in Azure ML SDK v2. By creating a data asset with a path pointing to this URI and then using pd.read_csv(my_path), the SDK handles the authentication and data access seamlessly, loading the CSV into a Pandas DataFrame.

  6. Question 6Intermediate

    Optimize language models for AI applications · Prepare for model optimization

    You are comparing two foundation models from the Azure AI Model Catalog, Model-A and Model-B, for a fine-tuning task. Model-A is larger and more capable but has higher inference costs. Model-B is smaller, faster, and cheaper but less powerful. You need to decide which model to fine-tune. What is the most cost-effective and technically sound approach to make this decision?

    Show answer & explanation

    Correct answer: C

    This approach balances cost and performance. Starting with the smaller, cheaper model is a cost-effective experiment. If Model-B can be fine-tuned to meet the required performance benchmarks, there is no need to incur the higher costs associated with the larger Model-A. This iterative approach, starting with the simplest/cheapest viable option, is a common best practice in machine learning.

  7. Question 7Intermediate

    Train and deploy models · Run model training scripts

    A data scientist has written a Python script to train a regression model. To ensure reproducibility, they have created a custom environment defined in a YAML file. The training script must be executed as a command job on a remote compute cluster. What is the correct way to specify the custom environment when defining the CommandJob using the Azure Machine Learning SDK v2?

    Show answer & explanation

    Correct answer: A

    In the Azure ML SDK v2, you can directly reference a local YAML file that defines an environment. When you construct the CommandJob, you set the environment parameter to the path of this file (e.g., environment='./my-env.yml'). Azure ML will then build or retrieve the specified environment and use it to run the job.

  8. Question 8Intermediate

    Design and prepare a machine learning solution · Determine the compute specifications for machine learning workload

    You are tasked with selecting a compute target for a model training job that involves a very large dataset (5 TB) and requires distributed training using Horovod. The training process is expected to run for several days. Which compute target offers the best performance and scalability for this specific workload?

    Show answer & explanation

    Correct answer: B

    An Azure Machine Learning Compute Cluster is the ideal choice for scalable, distributed training. It allows you to create a multi-node cluster that can automatically scale up and down. For distributed frameworks like Horovod, a multi-node GPU cluster provides the necessary parallel processing power and interconnectivity to efficiently train models on massive datasets.

  9. Question 9Intermediate

    Explore data, and run experiments · Use automated machine learning to explore optimal models

    A data scientist is using Automated ML for a time-series forecasting task to predict weekly sales. The dataset contains several years of historical data. To improve model accuracy, they want to incorporate features based on the week of the year and the day of the week. How can they ensure AutoML automatically generates these time-based features?

    Show answer & explanation

    Correct answer: B

    For forecasting tasks, AutoML has powerful built-in time-series featurization capabilities. By correctly identifying the time column in the forecasting settings and leaving featurization on 'auto' (the default), AutoML will automatically generate a rich set of features from the date/time column, such as year, month, day, day of week, week of year, and more.

  10. Question 10Intermediate

    Optimize language models for AI applications · Optimize through Retrieval Augmented Generation (RAG)

    You are building a RAG solution and have configured an Azure AI Search index as your vector store. You want to improve the relevance of search results by having the system understand the user's intent rather than just matching keywords. Which Azure AI Search feature should you enable and configure?

    Show answer & explanation

    Correct answer: C

    The semantic ranker (also known as semantic search) is a premium feature in Azure AI Search that uses deep learning models to understand the contextual meaning and intent behind a query. It re-ranks the initial set of results from keyword or vector search based on semantic relevance, significantly improving the quality of results for RAG applications.

Ready for the real thing?

The full DP-100 simulator has every exam-style question, timed mode, and instant scoring.

Go to the DP-100 simulator →