PDE Practice Exam: Professional Data Engineer
BigQuery, Dataflow, Pub/Sub, Dataproc — build GCP data pipelines that survive production.
What you'll be tested on
- Data Processing Systems
- Data Storage
- Data Analysis and ML
- Data Pipeline Orchestration
- Data Quality and Governance
Sample PDE questions
Your company built a TensorFlow neutral-network model with a large number of neurons and layers. The model fits well for the training data. However, when tested against new data, it performs poorly. What method can you employ to address this?
- Threading
- Serialization
- Dropout Methods
- Dimensionality Reduction
Show answer
C — Dropout MethodsThe symptoms describe overfitting: the model memorizes the training data but generalizes poorly to new data. Dropout is a regularization technique for neural networks that randomly disables a fraction of neurons during each training pass, preventing the network from becoming overly dependent on any single neuron and improving generalization. Threading and serialization are programming concerns unrelated to model accuracy. Dimensionality reduction can help with feature noise, but the standard exam answer for a large overfit neural network is dropout.
You are building a model to make clothing recommendations. You know a user's fashion preference is likely to change over time, so you build a data pipeline to stream new data back to the model as it becomes available. How should you use this data to train the model?
- Continuously retrain the model on just the new data.
- Continuously retrain the model on a combination of existing data and the new data.
- Train on the existing data while using the new data as your test set.
- Train on the new data while using the existing data as your test set.
Show answer
B — Continuously retrain the model on a combination of existing data and the new data.Because user preferences drift over time, the model must be continuously retrained, but training only on new data causes catastrophic forgetting of older patterns, so combining existing data with new data is correct. Using the new data solely as a test set never adapts the model to changing tastes, and training only on new data while testing on stale historical data mismeasures performance. Continuous retraining on the combined dataset keeps the recommendation model current while preserving knowledge from historical behavior, which is the standard pattern for streaming retraining pipelines.
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics. Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded. The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?
- Add capacity (memory and disk space) to the database server by the order of 200.
- Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.
- Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
- Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
Show answer
C — Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.The root problem is the single-table design with self-joins, which scales poorly as data volume grows 100-fold. Normalizing the schema into separate patient and visits tables eliminates expensive self-joins and lets the database engine use proper indexes and join strategies, fixing the performance and resource errors. Simply adding 200x capacity is costly and does not fix the design flaw. Sharding by date or partitioning by clinic are workarounds that complicate queries and still leave the inefficient self-join structure in place.
Access plans
| Access | Price |
|---|---|
| 3 months | |
| 1 year | |
| Lifetime |
Free preview inside — try 5 questions before you pay anything.
FAQ
How many practice questions are in this PDE bank?
349 questions covering the current PDE Professional Data Engineer syllabus, every one with the correct answer and an explanation.How long is the real PDE exam?
The official PDE exam gives you 120 minutes. Our timed exam mode uses the same limit so the pace feels familiar.What does PDE access cost?
Plans start at $3.99 for 3 months. One payment, no subscription — and far cheaper than retaking the real exam.