ADP Sample Questions & Answers
Free Associate Data Practitioner practice questions with worked answers and explanations. See how the ExamJungle simulator prepares you — then jump into the full test.
Launch the full ADP simulator →Showing 10 of 20 free samples.
- Question 1Intermediate
Data Management · Apply security measures and ensure compliance
True or False: Using a customer-managed encryption key (CMEK) with Cloud Storage means that Google no longer holds any component of the encryption key, and all cryptographic operations happen outside of Google Cloud.
Show answer & explanation
Correct answer: B
This statement is false. With CMEK, the customer manages the key within Google's Cloud Key Management Service (KMS). Google services (like Cloud Storage) still perform the cryptographic operations by requesting the KMS to use the key for encryption/decryption. The customer controls the key's lifecycle and access policies, but the key material resides within KMS. The scenario described (key held and used outside Google) applies to Customer-Supplied Encryption Keys (CSEK) or Cloud External Key Manager (EKM).
- Question 2Beginner
Data Preparation and Ingestion · Data quality and cleaning
A healthcare organization is building a data pipeline to process patient records. The raw data arrives as JSON files in a Cloud Storage bucket. The pipeline must de-identify sensitive information like patient names and social security numbers by applying masking transformations before loading the data into BigQuery for analysis. The solution must be a fully managed, graphical, low-code service to accelerate development. The pipeline should be designed according to the following flow:
graph TD A[GCS Bucket: Raw JSON] --> B{De-identification Pipeline}; B --> C[BigQuery Table: Anonymized Data];Which service should be used to build the de-identification pipeline (B)?
Show answer & explanation
Correct answer: C
Cloud Data Fusion is the ideal service for this requirement. It is a fully managed, cloud-native data integration service that provides a graphical interface and a broad library of pre-built transformations, including those for data masking and de-identification. This aligns perfectly with the 'low-code' and 'graphical' requirements. Dataflow would require custom coding, and Cloud Composer is an orchestrator, not a data transformation tool itself.
- Question 3Intermediate
Data Analysis and Presentation · Define, train, evaluate, and use ML models
You are building a regression model in BigQuery ML to predict housing prices. After training your model, you use the
ML.EVALUATEfunction and get the following output:{'mean_absolute_error': 25000, 'r2_score': 0.85}. What do these metrics signify about your model's performance?Show answer & explanation
Correct answer: C
For a regression model,
mean_absolute_error(MAE) indicates the average absolute difference between the predicted values and the actual values. An MAE of 25000 means the predictions are off by an average of $25,000. Ther2_score(R-squared) represents the proportion of the variance in the dependent variable (housing price) that is predictable from the independent variables. An r2_score of 0.85 means that the model's features explain 85% of the variability in housing prices, which is generally considered a good fit. - Question 4Beginner
Data Pipeline Orchestration · Schedule, automate, and monitor basic data processing tasks
An e-commerce company has a daily batch pipeline that updates product inventory in BigQuery. The pipeline is orchestrated by a Cloud Composer DAG. Recently, the DAG has been failing intermittently. You need to investigate the failures by reviewing the execution history, task logs, and the overall structure of the DAG. Which user interface should you use to perform this troubleshooting?
Show answer & explanation
Correct answer: C
Cloud Composer is a managed Apache Airflow service. The primary interface for managing, monitoring, and troubleshooting DAGs (Directed Acyclic Graphs) is the Airflow web UI. This interface provides detailed views of DAG runs, task statuses, logs for individual tasks, trigger history, and the rendered code. While logs are also available in Cloud Logging and the environment can be managed in the Google Cloud Console, the Airflow UI is the specific tool designed for in-depth DAG troubleshooting.
- Question 5Intermediate
Data Management · Configure access control and governance
A new data analyst has joined your team and needs permissions to run queries on all tables within the
production_analyticsdataset in BigQuery. They also need to be able to create new tables in this dataset. However, they must not be able to delete the dataset or modify its permissions. Following the principle of least privilege, which single predefined IAM role should you grant to the analyst at the dataset level?Show answer & explanation
Correct answer: B
The
BigQuery Data Editorrole is the most appropriate choice. It grants permissions to read, query, create, update, and delete tables within a dataset. This meets the requirements of running queries and creating new tables. It does not, however, grant permissions to delete the dataset itself or change its IAM policies, which are part of theBigQuery Data Ownerrole.BigQuery Data Vieweris too restrictive as it only allows reading data, not creating tables. - Question 6Intermediate
Data Management · Configure lifecycle management
A media company stores large video files in a Multi-Regional Cloud Storage bucket with Standard storage class. These files are accessed frequently for the first 30 days. After 30 days, access becomes rare, but the files must be kept for five years for compliance. After five years, they should be deleted. You need to implement a cost-effective lifecycle management policy. What should the lifecycle rule contain?
Show answer & explanation
Correct answer: B
This configuration is the most cost-effective. After the initial 30 days of frequent access, the data becomes cold. Transitioning it directly to Archive storage, which is designed for long-term, infrequent access, will provide the lowest storage cost for the five-year retention period. A final action to delete the objects after 1825 days (365 * 5) fulfills the compliance requirement. Using Nearline or Coldline would be more expensive than Archive for this long-term storage use case.
- Question 7Intermediate
Data Analysis and Presentation · Identify data trends, patterns, and insights by using BigQuery and Jupyter notebooks
You are analyzing a dataset of customer transactions in a Colab Enterprise notebook. The data is stored in a BigQuery table. You need to identify the top 5 customers by total spending. The BigQuery table is named
project.dataset.transactionsand has columnscustomer_idandpurchase_amount. You have already authenticated and initialized the BigQuery client in your notebook. Which Python code snippet using the BigQuery client library should you use?Show answer & explanation
Correct answer: A
This is the correct and standard way to execute a SQL query against BigQuery from a Python environment using the official client library. The SQL query correctly calculates the sum of
purchase_amountfor eachcustomer_id, orders the results in descending order to find the top spenders, and limits the output to the top 5. Theclient.query(sql).to_dataframe()method executes the query and conveniently converts the results into a Pandas DataFrame for further analysis in the notebook. - Question 8Intermediate
Data Pipeline Orchestration · Identify use cases for event-driven data ingestion from Pub/Sub to BigQuery
A startup is building an event-driven data processing system. When a new JSON file is uploaded to a Cloud Storage bucket, a lightweight data validation check needs to be performed. If the file is valid, a message should be sent to a Pub/Sub topic for downstream processing. The solution must be serverless, cost-effective for infrequent uploads, and require minimal infrastructure management. Which combination of services should be used to build this system?
sequenceDiagram participant User participant GCS as Cloud Storage participant Trigger participant Processor participant PubSub as Pub/Sub User->>GCS: Upload JSON file GCS->>Trigger: Event: object.finalize Trigger->>Processor: Invoke with event data Processor->>Processor: Validate JSON Processor->>PubSub: Publish messageShow answer & explanation
Correct answer: B
This is the classic serverless pattern for lightweight, event-driven tasks on Google Cloud. A Cloud Function can be directly triggered by a Cloud Storage event (like a file upload). The function can execute the validation logic and then publish to Pub/Sub. This solution is fully serverless, scales to zero (making it cost-effective for infrequent events), and requires no infrastructure management. Dataflow is better suited for large-scale stream or batch processing, not single-file validation. Cloud Run is also a good serverless option but Cloud Functions are often simpler for direct event handling like this.
- Question 9Advanced
Data Preparation and Ingestion · Data quality and cleaning
During a data quality audit, you discover that a critical
customerstable in BigQuery contains duplicate rows based on thecustomer_idcolumn. You need to create a new table,customers_deduped, that contains only the unique customer records. For duplicates, you must keep the record that was most recently updated, based on thelast_modified_tstimestamp column. Which SQL query should you use?Show answer & explanation
Correct answer: B
This query uses a window function
ROW_NUMBER()to rank rows for eachcustomer_idbased on thelast_modified_tsin descending order. This assigns a rank of 1 to the most recent record for each customer. TheQUALIFYclause, specific to BigQuery and some other data warehouses, filters the results of the window function, keeping only the rows where the rank is 1. This is an efficient and standard pattern for deduplication in BigQuery. The other options use incorrect SQL syntax or logic for this task. - Question 10IntermediateSelect 2
Data Management · Identify high availability and disaster recovery strategies
A gaming company is designing a high-availability and disaster recovery (HA/DR) strategy for its player profile data, which is stored in Cloud SQL for PostgreSQL. The primary business requirements are to ensure data durability against regional failures and to provide a read-only endpoint for analytics that does not impact the primary instance's performance. Which Cloud SQL features should be configured to meet these requirements? (Select TWO)
Show answer & explanation
Correct answers: A, B
Ready for the real thing?
The full ADP simulator has every exam-style question, timed mode, and instant scoring.