NCA-GENL Sample Questions

NCA-GENL Sample Questions & Answers

Walks through transformer architecture and machine learning basics, the largest block, deploying and integrating LLMs with prompt engineering, fine-tuning and evaluating models, prepping data for visualization, and weighing AI safety and ethics alongside governance.

Launch the full NCA-GENL simulator →

Showing 10 of 20 free samples.

  1. Question 1Intermediate

    Experimentation · Training Configuration

    A research team is fine-tuning a 70-billion parameter model on a single DGX node with 8 GPUs. The full model requires more VRAM than is available on a single GPU. To overcome this, they decide to split the model's layers across the 8 GPUs. What is this distributed training technique called?

    Show answer & explanation

    Correct answer: C

    Pipeline Parallelism is the technique of splitting the layers of a neural network model across multiple GPUs. Each GPU processes a subset of the model's layers (a stage) in a pipeline fashion. This is used when a model is too large to fit into a single GPU's memory. Data Parallelism replicates the model on each GPU, and Tensor Parallelism splits individual layers/tensors across GPUs.

  2. Question 2Beginner

    Experimentation · Evaluation Metrics

    When evaluating a text summarization model, a team calculates a score based on the overlap of n-grams between the machine-generated summary and a human-written reference summary. This metric is known as:

    Show answer & explanation

    Correct answer: C

    ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a set of metrics used for evaluating automatic summarization and machine translation. It works by comparing an automatically produced summary against a set of reference summaries (typically human-written) and measures the overlap of n-grams, word sequences, and word pairs.

  3. Question 3Intermediate

    Trustworthy AI · Security and Guardrails

    A developer is using the NVIDIA NeMo Framework to create a custom conversational AI application. They need to define rules for how the AI should respond to inappropriate user queries and ensure the conversation stays on a specific topic. Which NeMo component is specifically designed for this purpose?

    Show answer & explanation

    Correct answer: B

    NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable, rule-based controls to LLM-based applications. It allows developers to define specific conversational paths, control topics, prevent the model from accessing certain tools, and filter out undesirable content, making it the correct choice for ensuring safe and on-topic conversations.

  4. Question 4Beginner

    Core Machine Learning and AI Knowledge · Transformer Fundamentals

    What is the primary function of the self-attention mechanism in the Transformer architecture?

    Show answer & explanation

    Correct answer: C

    The self-attention mechanism allows the model to associate each word in the input sequence with other words in the same sequence. It calculates attention scores that determine how much focus to place on other words when encoding a specific word. This allows the model to capture long-range dependencies and understand context within the sequence, which is a key advantage over recurrent architectures.

  5. Question 5Intermediate

    Trustworthy AI · Responsible AI Practices

    A financial firm is using a generative AI model to create market analysis reports. They are concerned that the model, trained on public data, might inadvertently generate text that is too similar to copyrighted articles, creating a legal risk. Which AI safety problem does this scenario describe?

    Show answer & explanation

    Correct answer: C

    This scenario describes output regeneration, also known as plagiarism or regurgitation. It occurs when a generative model reproduces verbatim or near-verbatim excerpts from its training data. If the training data includes copyrighted material, this can lead to copyright infringement. This is a key concern in responsible AI development.

  6. Question 6Intermediate

    Software Development · Model Deployment with NVIDIA Tools

    An AI engineer is optimizing a deployed Mistral-7B model using TensorRT-LLM. They want to reduce memory footprint and increase inference speed by using lower-precision numerical formats for model weights, with minimal impact on accuracy. Which TensorRT-LLM feature is most appropriate for this goal?

    Show answer & explanation

    Correct answer: B

    Weight-only quantization is a technique used to represent a model's weights using lower-precision integers (like INT8 or INT4) instead of floating-point numbers (like FP16). This significantly reduces the model's memory footprint and can accelerate inference, especially on GPUs with specialized hardware for integer arithmetic. TensorRT-LLM provides robust support for these quantization methods.

  7. Question 7Beginner

    Software Development · Prompt Design Techniques

    A developer is crafting a prompt for an LLM to generate a Python function that sorts a list of dictionaries by a specific key. To improve the model's reasoning process and increase the likelihood of a correct output, they decide to add the instruction 'Think step by step'. This is an example of what type of prompting technique?

    Show answer & explanation

    Correct answer: C

    Chain-of-thought (CoT) prompting is a technique that encourages the LLM to break down a complex problem into intermediate reasoning steps before providing a final answer. The phrase 'Think step by step' is a classic way to elicit this behavior, often leading to more accurate and logical results, especially for arithmetic, commonsense, and symbolic reasoning tasks.

  8. Question 8Beginner

    Core Machine Learning and AI Knowledge · Neural Network Basics

    Which of the following activation functions is commonly used in the output layer of a neural network for multi-class classification problems to convert logits into probability distributions?

    Show answer & explanation

    Correct answer: D

    The Softmax function is used in the output layer for multi-class classification. It takes a vector of real-valued scores (logits) and transforms them into a probability distribution, where each value is between 0 and 1, and the sum of all values equals 1. This makes it ideal for representing the probability of an input belonging to each of the possible classes.

  9. Question 9Intermediate

    Experimentation · Fine-tuning Strategies

    A startup is building a customer support chatbot. They have limited GPU resources and cannot afford to fully fine-tune a large language model. They need a method to adapt a pre-trained model to their company's specific support documents and tone. Which of the following is the MOST resource-efficient fine-tuning strategy?

    Show answer & explanation

    Correct answer: B

    Parameter-Efficient Fine-Tuning (PEFT) methods, such as LoRA (Low-Rank Adaptation), are designed for this exact scenario. They freeze the vast majority of the pre-trained model's parameters and only train a small number of additional parameters (adapters). This dramatically reduces memory and compute requirements, making it possible to fine-tune large models on consumer-grade or limited enterprise hardware.

  10. Question 10Advanced

    Software Development · Model Deployment with NVIDIA Tools

    An organization is deploying a generative AI application using the recently announced NVIDIA NIM (NVIDIA Inference Microservices). What is the primary advantage of using NIM for deployment?

    Show answer & explanation

    Correct answer: C

    NVIDIA NIM (NVIDIA Inference Microservices) provides pre-built, cloud-native microservices that simplify the deployment of generative AI models. It packages models and their dependencies into optimized containers, exposing a standard industry API. This abstracts away the complexity of the underlying inference stack (like TensorRT-LLM and Triton), allowing developers to deploy highly performant models much more quickly and easily across various environments.

Ready for the real thing?

The full NCA-GENL simulator has every exam-style question, timed mode, and instant scoring.