NCP-AIO Sample Questions

NCP-AIO Sample Questions & Answers

Tests your grasp of setting up Mission Control, Base Command Manager and job scheduling, the single biggest piece, plus cluster administration, GPU virtualization, deploying training and inference workloads, and chasing down container problems.

Launch the full NCP-AIO simulator →

Showing 9 of 19 free samples.

  1. Question 1Intermediate

    Administration · MIG Configuration

    A hospital's data science team uses a shared NVIDIA DGX H100 server for developing medical imaging AI models. To ensure resource isolation and predictable performance for multiple concurrent users, the lead AI Operations engineer decides to partition the GPUs using Multi-Instance GPU (MIG). The requirement is to create the maximum possible number of isolated, compute-capable instances on a single H100 GPU. Which nvidia-smi command correctly achieves this?

    Show answer & explanation

    Correct answer: A

    The NVIDIA H100 GPU supports up to seven MIG instances. The smallest compute-capable profile is 1g.10gb. This command correctly creates seven instances of this profile, maximizing the number of isolated environments on a single GPU. The other options either use invalid profile names for H100, an incorrect number of instances, or a mix of profiles that does not yield the maximum number.

  2. Question 2Intermediate

    Administration · Data Center Architecture

    An AI infrastructure team is managing a large cluster with hundreds of nodes. They need a way to perform out-of-band management, including power cycling, BIOS configuration, and monitoring hardware sensors, without relying on the node's primary operating system. The cluster nodes are equipped with NVIDIA BlueField-3 DPUs. How can the team achieve this level of remote management?

    Show answer & explanation

    Correct answer: B

    NVIDIA BlueField-3 DPUs include an integrated Baseboard Management Controller (BMC). This allows for true out-of-band management of the host server. Administrators can connect to the DPU's management port to perform tasks like power control, firmware updates, and sensor monitoring, completely independent of the host server's OS state.

  3. Question 3Intermediate

    Troubleshooting and Optimization · Network Performance

    A system administrator is troubleshooting a performance issue with a distributed training job running across several nodes connected via an InfiniBand fabric. They suspect a faulty cable or switch port is causing excessive errors on a specific link. Which command-line utility is the MOST appropriate first step to check the error counters and status of all links in the InfiniBand fabric from a single node?

    Show answer & explanation

    Correct answer: D

    ibdiagnet is a comprehensive InfiniBand fabric diagnostic tool. It scans the entire fabric topology, checks for connectivity, and reports on the health of links, including detailed error counters like symbol errors and link recovery events. Running ibdiagnet is the standard procedure for identifying physical layer issues such as bad cables or failing ports.

  4. Question 4Advanced

    Installation and Deployment · Slurm Installation and Configuration

    An administrator is configuring a new Slurm cluster with NVIDIA H100 GPUs. They have enabled MIG on all GPUs and created various MIG instance profiles. The goal is to allow users to request a specific MIG instance profile in their sbatch scripts. For example, a user should be able to request two instances of the 2g.20gb profile. How must the gres.conf file be configured on the Slurm controller to enable this functionality?

    Show answer & explanation

    Correct answer: D

    To make Slurm aware of MIG instances and allow users to request them by profile, you must use the MIG_CONFIG flag for the GPU resources in gres.conf. This tells Slurm to query the MIG configuration for each GPU and make the available profiles (e.g., 2g.20gb, 3g.40gb) available as schedulable resources. Users can then request them using --gpus=mig:2g.20gb:2.

  5. Question 5Advanced

    Administration · vGPU Administration

    A financial services company is using NVIDIA AI Enterprise to run critical, low-latency inference workloads on a vSphere cluster. During a routine audit, a security team member raises a concern that a compromised VM could potentially access the memory of other VMs running on the same GPU via a side-channel attack. Which NVIDIA vGPU feature, when enabled on the host, mitigates this specific security risk?

    Show answer & explanation

    Correct answer: C

    Frame buffer (VRAM) scrubbing is a security feature of NVIDIA vGPU. When enabled, it ensures that whenever a VM's vGPU instance is destroyed or its memory is deallocated, the corresponding physical VRAM region on the GPU is overwritten with a pattern (typically zeros). This prevents a subsequently scheduled VM from potentially reading residual data left by the previous VM, mitigating the risk of data leakage between tenants.

  6. Question 6Beginner

    Installation and Deployment · Mission Control Toolkit

    An administrator is deploying a new AI cluster using the NVIDIA Mission Control toolkit. The process involves several stages, from hardware validation to software stack deployment. The following diagram shows a simplified workflow. At which stage would the administrator use the toolkit's capabilities to perform stress tests and benchmarks on GPU, network, and storage components to establish a performance baseline?

    flowchart TD A[Start] --> B(Hardware Discovery); B --> C{Firmware & BIOS Update}; C --> D(Bare-metal OS Provisioning); D --> E(System Validation & Benchmarking); E --> F(Deploy Schedulers - Slurm/K8s); F --> G[End: Cluster Ready];
    Show answer & explanation

    Correct answer: D

    The NVIDIA Mission Control toolkit is designed to streamline cluster deployment. After the base OS and drivers are provisioned (Stage D), the next critical step (Stage E) is to validate that all hardware components are functioning correctly and performing as expected. This involves running a suite of tests, such as the High-Performance Linpack (HPL) benchmark, GPU burn-in tests, and network/storage I/O tests, to establish a performance baseline before the cluster is put into production.

  7. Question 7Advanced

    Troubleshooting and Optimization · Magnum IO Components

    A large-scale language model (LLM) training job, which uses GPUDirect Storage (GDS), is experiencing I/O bottlenecks. The cluster uses a Lustre parallel file system. A senior engineer suspects that the issue is related to how data is being transferred between the storage and GPU memory. Which of the following configurations is required to ensure data is transferred directly from the Lustre file system to the GPU memory, bypassing the CPU and system RAM?

    Show answer & explanation

    Correct answer: B

    For GPUDirect Storage to function correctly, several components must be in place. Critically, the nvidia-peermem module is required to facilitate direct memory access between the GPU and third-party devices like NVMe drives. Furthermore, the file system itself (in this case, Lustre) must be mounted with a specific option (e.g., gds) to signal that it is GDS-aware and should use the direct data path when possible. Without both of these, the data path will likely fall back to the conventional route through the CPU's main memory.

  8. Question 8Intermediate

    Installation and Deployment · BCM Deployment and Configuration

    You are the lead operations engineer for an AI startup that has just received a new multi-node server with NVIDIA Blackwell GPUs. You are tasked with setting up the initial software environment using Base Command Manager (BCM). You need to deploy a consistent OS image, NVIDIA drivers, CUDA toolkit, and DCGM to all nodes. What is the correct sequence of high-level steps to accomplish this using BCM?

    Show answer & explanation

    Correct answer: C

    The standard BCM workflow for initial provisioning is to first discover the nodes via their out-of-band BMC interfaces. Once discovered, they are organized into a logical node group. Next, a comprehensive software image is defined, which includes not just the base OS but also the entire software stack (drivers, CUDA, DCGM, etc.). Finally, this software image is assigned to the node group, and BCM orchestrates the automated, parallel provisioning process across all nodes.

  9. Question 9Intermediate

    Administration · Run:ai Administration

    A data science team reports that their JupyterLab environment, running as a pod in a Kubernetes cluster, is frequently being terminated. The cluster uses Run:ai for workload scheduling. An investigation of the Run:ai dashboard shows the pod was 'preempted'. What is the most likely reason for this preemption?

    Show answer & explanation

    Correct answer: C

    Run:ai enhances Kubernetes with advanced scheduling features, including preemption and fair-share policies. When a higher-priority job is submitted and there are no free resources, the Run:ai scheduler can preempt (pause or terminate) a lower-priority interactive job, like a JupyterLab session, to free up the necessary GPUs. The preempted job is typically requeued and will resume once resources become available again.

Ready for the real thing?

The full NCP-AIO simulator has every exam-style question, timed mode, and instant scoring.