PTCA Sample Questions & Answers
Core concepts, tensors and training and testing models carry the most weight, ahead of performance topics like precision and distributed training, building neural network blocks, and handling data with datasets, DataLoaders and transforms.
Launch the full PTCA simulator →Showing 6 of 12 free samples.
- Question 1Intermediate
PyTorch Fundamentals · Device placement and management
During the setup phase of a deep learning pipeline, an architect configures the model and optimizer. Which of the following sequences represents the correct best practice for initializing an optimizer in relation to moving the model to a target hardware device (e.g., GPU)?
Show answer & explanation
Correct answer: B
You must move the model to the target device before passing its parameters to the optimizer. If you initialize the optimizer first and then move the model, the optimizer will hold references to the old CPU parameters, while the model will use the newly allocated GPU parameters, causing the optimizer step to have no effect on the model's actual GPU weights.
- Question 2Advanced
PyTorch Fundamentals · Tensor creation and operations
A computer vision model requires blending a feature map tensor
Awith a per-channel bias tensorB. TensorAhas the shape(16, 3, 256, 256)representing (batch, channels, height, width). TensorBhas the shape(3, 1, 1). When the operationC = A + Bis executed, what will be the resulting shape of tensorCbased on PyTorch broadcasting semantics?Show answer & explanation
Correct answer: A
According to PyTorch broadcasting rules, dimensions are aligned from right to left. B's shape (3, 1, 1) aligns with the last three dimensions of A (3, 256, 256). Since dimensions of size 1 can be broadcast to match the other tensor, and missing leading dimensions (the batch dimension 16) are implicitly assumed to be 1, the operation broadcasts successfully resulting in the maximum size along each dimension: (16, 3, 256, 256).
flowchart LR A["Tensor A: (16, 3, 256, 256)"] --> Ops(+) B["Tensor B: ( 1, 3, 1, 1)"] --> Ops Ops --> C["Result C: (16, 3, 256, 256)"] - Question 3Beginner
PyTorch Fundamentals · PyTorch core concepts and workflow
To compute the gradients of the loss with respect to all tensors in the computation graph that have requires_grad=True, you must call the ________ method on the final loss tensor.
Show answer & explanation
Correct answer: A
The backward() method is called on the scalar loss tensor to traverse the computation graph backwards, computing the gradient of the loss with respect to all leaf tensors that have requires_grad=True.
- Question 4Advanced
PyTorch Fundamentals · Training / evaluation / inference workflow
A data scientist observes that their model's training loss is behaving erratically, oscillating wildly and occasionally exploding to infinity, despite a very small learning rate. They review their core training loop:
for data, target in dataloader: output = model(data) loss = criterion(output, target) loss.backward() optimizer.step()What critical omission in this training loop is causing the erratic loss behavior?
Show answer & explanation
Correct answer: C
By default, PyTorch accumulates gradients in the
.gradattributes of tensors duringloss.backward(). Ifoptimizer.zero_grad()is not called before the backward pass, the gradients from the current batch are added to the gradients from all previous batches. This quickly leads to massive gradient values and unstable, exploding loss.flowchart TD A[Forward Pass] --> B[Compute Loss] B --> C{zero_grad() called?} C -->|Yes| D[backward: Gradients=Current] C -->|No| E[backward: Gradients=Current + Old] D --> F[optimizer.step() - Stable] E --> G[optimizer.step() - Explodes!] - Question 5Intermediate
PyTorch Fundamentals · Tensor creation and operations
When attempting to alter the shape of a tensor
xusingx.view(-1, 128), a developer encounters aRuntimeErrorstating that the tensor is not contiguous in memory. Which alternative method should they use to guarantee the shape change will succeed regardless of the tensor's memory layout?Show answer & explanation
Correct answer: B
The .reshape() method is the robust alternative to .view(). While .view() strictly requires the underlying memory to be contiguous, .reshape() will first check if the tensor is contiguous. If it is, it returns a view; if it is not, it automatically copies the data to a contiguous block of memory and then returns the view. This guarantees the operation succeeds.
- Question 6IntermediateSelect 2
PyTorch Fundamentals · Device placement and management
A developer is writing a training script that must run natively with hardware acceleration on modern Apple Silicon Macs. Which TWO of the following statements correctly describe how to implement device management for this environment in PyTorch? (Select TWO)
Show answer & explanation
Correct answers: A, B
To utilize Apple Silicon GPUs, PyTorch provides the MPS (Metal Performance Shaders) backend. You verify its presence with
torch.backends.mps.is_available()and instantiate the device usingtorch.device('mps'). CUDA is strictly for NVIDIA hardware.To utilize Apple Silicon GPUs, PyTorch provides the MPS (Metal Performance Shaders) backend. You verify its presence with
torch.backends.mps.is_available()and instantiate the device usingtorch.device('mps'). CUDA is strictly for NVIDIA hardware.
Ready for the real thing?
The full PTCA simulator has every exam-style question, timed mode, and instant scoring.