Generative AI and Large Language Models Questions
The capabilities and behavior of modern generative and large language models. Covers how LLMs are pretrained, in-context learning and few-shot prompting, generative model families (autoregressive, diffusion), context windows, and tokenization and sampling. Emphasizes understanding what generative models can and cannot do and how they differ from discriminative ML.
In RLHF, what is the purpose of applying a KL penalty relative to a reference policy? Explain how it guards against extreme policy shifts, preserves pre-trained behavior, and how you might tune the KL coefficient in practice.
Sample Answer
Direct answer: The KL (Kullback-Leibler divergence) penalty in RLHF measures how far the current policy has drifted from a fixed reference policy (usually the supervised-fine-tuned starting point) and subtracts a term proportional to that divergence from the training objective, which keeps the policy from making large, uncontrolled changes purely to chase reward.
Structured elaboration: Without a KL penalty, a policy-optimization algorithm like PPO can, in principle, keep increasing reward by drifting arbitrarily far from reasonable, fluent language, especially if the reward model has any exploitable blind spot, since nothing in the objective directly penalizes that drift. The KL penalty term, typically reward−βDKL(πθ∥πref), directly trades off reward against how different the new policy's output distribution is from the reference model's, so it preserves general language quality and previously-learned behavior even while the policy adapts to the reward signal. In practice the coefficient β is tuned adaptively in many implementations: if the measured KL divergence for a batch exceeds a target value, β is increased for the next update (pulling the policy back toward the reference), and if KL is comfortably below target, β is decreased slightly to allow more room for improvement.
Worked example: If a batch's rollouts show a KL divergence of 0.02 against a target of 0.01, a common adaptive scheme (roughly following the original PPO-KL adaptive-penalty approach) would increase β, for example multiplying it by 1.5, for the next update; if instead the observed KL were 0.002, well under target, β might be reduced, for example halved, to let the policy move more freely while it remains far from the drift limit.
Trade-offs and pitfalls: Setting β too high effectively freezes the policy close to its starting point, wasting the RLHF stage's ability to improve on the reward signal at all; setting it too low reintroduces the risk of unconstrained drift and reward hacking that the penalty exists to prevent. Because the reference policy is fixed throughout training, if the supervised-fine-tuned starting point itself has quality issues, the KL penalty will actively work against fixing them, since it specifically discourages moving away from that reference.
At a high level, describe diffusion generative models: what are the forward (noising) and reverse (denoising) processes, how is the model trained, and how does sampling work at generation time? Give an example use case where diffusion is preferred over GANs.
Sample Answer
Direct answer
Diffusion models generate data by learning to reverse a gradual noising process: a fixed forward process progressively adds Gaussian noise to real data over many steps until it becomes indistinguishable from pure noise, and the model learns the reverse process, predicting the noise (or an equivalent quantity) at each step so that starting from pure noise and iteratively denoising produces a realistic sample.
Structured elaboration
Forward (noising) process. Starting from a real data point x0, Gaussian noise is added in small increments over T steps according to a fixed noise schedule βt, producing progressively noisier versions x1,x2,…,xT, until xT is essentially pure noise with no remaining signal from x0. This process has no learned parameters; it's a fixed, known mathematical procedure.
Reverse (denoising) process. The model learns to approximate the reverse: given a noisy xt, predict either the noise that was added to produce it, or directly the mean of the slightly-less-noisy xt−1. At generation time, you start from pure noise xT and repeatedly apply the learned reverse step, T times, gradually removing noise until you arrive at a sample x0 that looks like it came from the real data distribution.
Training objective. The most common formulation trains the model to predict the noise ϵ that was added at a randomly sampled step t, using a simple mean-squared-error loss between the predicted and true noise. This is mathematically equivalent (up to a reweighting) to optimizing a variational lower bound on the data likelihood, but the simplified noise-prediction MSE objective is what's used in practice because it trains more stably and produces better samples than directly optimizing the full likelihood bound.
Why Gaussian noise and reweighted objectives. Gaussian noise is used because it has convenient closed-form properties (you can jump directly from x0 to any noisy xt in one step during training, without simulating all the intermediate steps, since the sum of the individually-added Gaussian noise steps is itself Gaussian). The training objective is typically reweighted across timesteps (rather than using the literal variational-bound weighting) because empirically it puts more training emphasis on the timesteps that matter most for final sample quality, which the raw likelihood bound underweights.
Diffusion versus GANs, an example use case. Diffusion is generally preferred over GANs when sample quality and diversity matter more than generation speed, and when training stability is a priority, e.g., a text-to-image product where users are willing to wait a few seconds for a high-quality, varied image, versus a real-time application (live video style transfer) where a GAN's single-forward-pass speed is the deciding factor even at some cost to quality or diversity.
Trade-offs & pitfalls
The biggest practical cost of the diffusion formulation is sampling speed: generating one sample requires running the learned reverse step many times (historically hundreds to a thousand steps, though modern fast samplers have brought this down substantially), which is fundamentally more expensive at inference time than a GAN's single forward pass. A common misunderstanding is treating the number of diffusion steps as purely a quality knob you can freely increase; beyond a point, more steps mostly add latency without meaningfully improving sample quality, and the practical tuning is finding the fewest steps that preserve acceptable quality for the specific application.
Explain the difference between prompt engineering and prompt tuning (parameter-efficient weight updates). Provide a scenario where prompt tuning is preferable to prompt engineering, and outline how you would evaluate prompt-tuned adapters against hand-crafted prompts.
Sample Answer
Direct answer: Prompt engineering designs the input text (templates, few-shot examples, phrasing) with no change to model parameters at all, while prompt tuning learns a small set of trainable "soft prompt" parameters (continuous embeddings, not real word tokens) while keeping the base model frozen, so it modifies behavior more reliably than hand-crafted text alone while still using far fewer resources than full fine-tuning.
Structured elaboration: Prompt engineering is fast, free of training cost, and easy to iterate on, but it is fundamentally limited by what can be expressed in natural-language instructions and few-shot examples, and it can be brittle across slightly different input styles or distributions, since nothing about the model itself has changed to accommodate the task. Prompt tuning trains a set of continuous vectors (conceptually like extra "virtual tokens") prepended to the input, optimized directly against labeled examples, so it can encode task-specific signal that a hand-written prompt cannot easily express in natural language, while still touching a tiny fraction of the model's total parameters and leaving the base model's weights untouched.
A scenario favoring prompt tuning: a production classification pipeline over domain-specific legal documents, where consistent, high accuracy and robustness matter, a moderate labeled set exists (on the order of 1,000 to 10,000 examples), and the deployment can accommodate a small amount of added parameters. Hand-crafted prompts tend to hit a ceiling here and be brittle across different document styles or drafting conventions, while a prompt-tuned adapter can absorb domain-specific patterns that manual prompt engineering struggles to capture.
Worked example: To evaluate prompt-tuned adapters against hand-crafted prompts rigorously: build train/validation/test splits that include both in-distribution and out-of-distribution document styles, compare against strong baselines (best hand-crafted prompt, few-shot prompting, zero-shot, and the prompt-tuned variant), measure not just task accuracy but calibration, robustness (the accuracy drop specifically on the OOD split), latency, and storage cost, and run each comparison across multiple random seeds with a paired significance test rather than a single run. Varying the labeled-data size (for example 10, 100, and 1,000 examples) while repeating this comparison reveals the DATA-EFFICIENCY trade-off directly: if prompt tuning only pulls ahead of hand-crafted prompting once several hundred labeled examples exist, that tells you exactly the labeled-data threshold at which it becomes worth the added infrastructure.
Trade-offs and pitfalls: Evaluating prompt tuning against hand-crafted prompts only on in-distribution accuracy misses the point of the comparison, since the more common real-world advantage of prompt tuning is ROBUSTNESS across document styles, not necessarily a large in-distribution accuracy gain, so the OOD split is not optional in this evaluation. Choosing prompt tuning purely because it is "more sophisticated" than prompt engineering, when a well-crafted prompt with a few good examples already meets the accuracy and robustness bar, adds deployment complexity (an adapter to store, version, and roll back) for no measurable benefit.
You must decide between two third-party LLM options for a knowledge assistant: a faster, cheaper model with slightly lower factual accuracy, versus a slower, costlier model with better factuality. How would you evaluate and choose, and how might you combine both to meet product goals?
Sample Answer
Direct answer
You must decide between two third-party LLMs for a knowledge assistant: a faster, cheaper model with slightly lower factual accuracy versus a slower, costlier model with better factuality, and possibly using both together. The right approach is to define a small set of measurable evaluation criteria tied to the actual product requirement, benchmark both models against real (or realistic) queries on those criteria, and then decide whether a single model suffices or whether a routing/fallback strategy combining both is worth the added complexity.
Structured elaboration
How to evaluate. Build a held-out evaluation set of realistic queries with known-correct answers (or human-graded rubrics for open-ended ones), and measure each candidate model on: factual accuracy (does it get the answer right, and does it hallucinate on out-of-scope questions), latency (p50/p95 response time under realistic load), cost per query at your expected volume, and any user-experience metrics that matter for the product (helpfulness ratings, task completion rate in A/B tests). Benchmarks alone are not sufficient; a model can score well on a public benchmark and still perform differently on your specific domain and query distribution, so the evaluation set must reflect your actual traffic.
Vendor/integration risk. Beyond raw model quality, evaluate SLA guarantees, rate limits, data-handling and privacy terms, and how exposed you are if the vendor changes pricing, deprecates the model, or has an outage, since a third-party dependency carries operational risk that a benchmark score alone won't capture.
Combining both models. A common production pattern is routing: use the fast, cheap model for the majority of queries (the ones it handles well), and escalate to the slower, more accurate model only for queries flagged as higher-risk or lower-confidence, e.g., via the fast model's own confidence signal, query complexity heuristics, or a lightweight classifier trained on where the fast model tends to fail. This captures most of the cost savings of the cheap model while limiting factual-accuracy risk to the harder subset of queries that actually need it.
Worked example
Say the cheap model costs $0.20 per 1,000 queries and the accurate model costs $2.00 per 1,000 queries, roughly a 10x cost difference, and evaluation shows the cheap model is factually correct 92% of the time versus 98% for the expensive model on your held-out set. If a routing classifier can reliably identify the roughly 20% of queries where the cheap model is most likely to be wrong (say, questions requiring precise numeric facts or recent information) and escalate only those to the expensive model, the blended cost is 0.8×$0.20+0.2×$2.00=$0.16+$0.40=$0.56 per 1,000 queries, about a 72% cost reduction from always using the expensive model, while capturing most of its accuracy benefit specifically where it matters most.
Trade-offs & pitfalls
A routing strategy only pays off if the signal for "this query needs the accurate model" is genuinely predictive; a poorly calibrated router either escalates too much (eroding the cost savings) or too little (letting factuality-sensitive queries slip through to the cheap model). It also adds real engineering and operational complexity: two vendor integrations, two SLAs to monitor, and a routing component that itself needs to be evaluated and maintained over time. For an early-stage product without the traffic volume or engineering capacity to build and maintain a router well, a single well-chosen model, even if slightly suboptimal on cost or accuracy, is often the more pragmatic starting point, with routing revisited once volume and evaluation infrastructure justify the added complexity.
What does it mean for an LLM-based system to act as an agent that calls external tools (search, calculator, code execution)? At a conceptual level, what is the basic risk of giving a model tool access, and what is one simple mitigation (e.g., scoping what a tool call can do)?
Sample Answer
Direct answer
An LLM acting as an agent that calls external tools means the model doesn't just generate text; it can decide, based on the conversation, to invoke an external function such as a search API, a calculator, or a code execution environment, and then incorporate that tool's real output into its next response. The basic risk of giving a model tool access is that the model's decision about when and how to call a tool is itself just another generated output, so a model that has been manipulated, is confused, or is simply wrong can call a tool with harmful, incorrect, or unintended arguments. A simple mitigation is scoping exactly what a given tool call is allowed to do, so that even a bad decision by the model cannot cause serious harm.
Structured elaboration
How tool calling actually works. The model is given a description of available tools, their names, purposes, and expected arguments, as part of its context. When the model decides a tool is needed, it generates a structured request, for example a function name and arguments, instead of, or as part of, its normal text output. That request is executed by the surrounding system, not by the model itself, and the tool's real result is fed back into the model's context so it can use that result in its next generated response. The core capability being exercised is still just token generation, applied to producing a structured "call this function with these arguments" output instead of ordinary prose.
The basic risk. Because the decision to call a tool, and with what arguments, is generated by the model the same way any other text is generated, it inherits every weakness of LLM generation: it can be wrong, it can be manipulated by adversarial or unexpected input (for example a user, or even a retrieved document, tricking the model into calling a tool it shouldn't), and it can be confidently incorrect about whether calling the tool is the right move at all. A tool call is more consequential than ordinary text generation specifically because it has a real-world side effect, running code, querying a database, sending a request, so a bad decision here can cause actual harm rather than merely producing a wrong sentence.
A basic mitigation: scoping. Rather than giving a model broad, unrestricted access to a powerful capability, for example "run arbitrary code" or "access any file," scope each tool narrowly to exactly what it needs to do, for example "look up a value in this specific read-only table" instead of "run any database query." That way, even if the model's decision to call the tool is wrong or manipulated, the blast radius of what that call can actually do is small and bounded by design, rather than depending entirely on the model always deciding correctly.
Trade-offs & pitfalls
Scoping tools narrowly is a real trade-off against flexibility: a narrowly scoped tool can do less, which sometimes means the agent genuinely cannot accomplish a legitimate task it otherwise could with broader access. The common mistake is treating "the model is capable, so it will call tools appropriately" as sufficient safety, when the model's tool-calling decisions carry the same reliability limitations as any other LLM output. The safety property has to come from constraining what a tool call CAN do, through scoping, permissions, and sandboxing, not from trusting the model to always decide correctly.
Unlock Full Question Bank
Get access to all 37 Generative AI and Large Language Models interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.