GCP-PDE Sample Questions & Answers
Planning, building, and operationalizing data pipelines carries the biggest weight, alongside designing for security, reliability, and portability, choosing storage systems and warehouses, preparing data for visualization or AI, and automating workloads.
Launch the full GCP-PDE simulator →Showing 6 of 12 free samples.
- Question 1Advanced
Storing the data · Selecting storage systems
You are designing a database schema for Cloud Bigtable to store time-series data from environmental sensors. The query patterns involve retrieving data for a specific sensor over a specific time range. You also need to perform periodic batch analytics on all sensors in a specific geographic region. The current row key design is
[SensorID]#[Timestamp].What potential issue does this design introduce, and how should you fix it?
Show answer & explanation
Correct answer: B
If SensorIDs are sequential or if specific sensors are much more active, writing sequentially can cause hotspotting on specific nodes. Hashing the SensorID distributes writes across the cluster. While
[SensorID]#[Timestamp]is good for single-sensor queries, it doesn't solve the write distribution issue if IDs are close. However, the best answer for general hotspot prevention in this context is hashing or salting. Note: For the geographic query requirement, a different approach might be needed (like a secondary index or different table), but the primary issue with[SensorID]is often write distribution if not hashed. - Question 2Intermediate
Storing the data · Planning for using a data warehouse
A retail company uses BigQuery for its data warehouse. They have a table named
Salesthat is 500 TB in size. Analysts frequently run queries filtering byTransactionDateand aggregating byStoreId. These queries are becoming slower and more expensive.How should you optimize the table schema to improve performance and reduce costs?
Show answer & explanation
Correct answer: A
Partitioning by
TransactionDateensures that queries filtering on date only scan the relevant partitions, drastically reducing cost (bytes scanned). Clustering byStoreIdsorts the data within each partition, which accelerates aggregations and filters on that column. This combination directly addresses both the filter and aggregation requirements. - Question 3Beginner
Preparing and using data for analysis · Preparing data for AI and ML
Your organization is building a machine learning pipeline to predict customer churn. The data resides in BigQuery, and your data science team wants to use SQL to build and train the model because they are less proficient in Python/TensorFlow. The model needs to be retrained weekly with fresh data.
Which Google Cloud service should you use?
Show answer & explanation
Correct answer: A
BigQuery ML allows users to create and execute machine learning models in BigQuery using standard SQL queries. This fits the requirement of using SQL and avoids moving data out of the warehouse. Vertex AI Custom Training would require Python/container knowledge.
- Question 4Intermediate
Maintaing and automating data workloads · Monitoring and troubleshooting processes
You are managing a Dataflow pipeline that processes streaming data from Pub/Sub. You notice that the pipeline's system lag is increasing effectively unbounded. Upon investigation, you see that the CPU utilization on the worker nodes is consistently near 100%.
Which action should you take to resolve this performance issue?
Show answer & explanation
Correct answer: A
High CPU utilization with increasing lag indicates the pipeline is under-provisioned for the volume of data or complexity of processing. Enabling Streaming Engine moves state storage and shuffle off the workers, reducing CPU load. Alternatively, increasing the max workers allows the autoscaler to provision more compute power. Switching to a larger machine type is also valid but requires a pipeline update/restart.
- Question 5Intermediate
Preparing and using data for analysis · Sharing data
Your company needs to share a 2 PB dataset residing in BigQuery with a partner organization. The partner also uses Google Cloud. You need to share the data securely without copying it, and the partner will pay for their own query costs.
Which solution should you implement?
Show answer & explanation
Correct answer: A
Analytics Hub is the modern, managed way to share assets in BigQuery. It allows you to publish datasets as listings. Subscribers (the partner) get a linked dataset in their project. They query the linked dataset, and the compute costs are billed to them, while storage remains with the provider. No data copying occurs. Authorized Views are also a valid mechanism but Analytics Hub provides better management for external sharing.
- Question 6Advanced
Maintaining and automating data workloads · Designing automation and repeatability
You are designing a data pipeline that uses Cloud Composer (Apache Airflow) to orchestrate a sequence of BigQuery jobs. One specific task waits for a file to arrive in a Cloud Storage bucket before triggering the next step. The file arrival time is unpredictable, ranging from 1 minute to 12 hours after the previous task.
Which Airflow operator or sensor strategy is best for cost and performance?
Show answer & explanation
Correct answer: A
Standard sensors occupy a worker slot while waiting (polling), which is expensive for long wait times (up to 12 hours). Deferrable operators (async) release the worker slot and offload the polling trigger to the Airflow Triggerer service, significantly reducing resource consumption and cost on the Composer cluster.
Ready for the real thing?
The full GCP-PDE simulator has every exam-style question, timed mode, and instant scoring.