Generative AI and Large Language Models Questions
The capabilities and behavior of modern generative and large language models. Covers how LLMs are pretrained, in-context learning and few-shot prompting, generative model families (autoregressive, diffusion), context windows, and tokenization and sampling. Emphasizes understanding what generative models can and cannot do and how they differ from discriminative ML.
You must decide between two third-party LLM options for a knowledge assistant: a faster, cheaper model with slightly lower factual accuracy, versus a slower, costlier model with better factuality. How would you evaluate and choose, and how might you combine both to meet product goals?
Sample Answer
Direct answer
You must decide between two third-party LLMs for a knowledge assistant: a faster, cheaper model with slightly lower factual accuracy versus a slower, costlier model with better factuality, and possibly using both together. The right approach is to define a small set of measurable evaluation criteria tied to the actual product requirement, benchmark both models against real (or realistic) queries on those criteria, and then decide whether a single model suffices or whether a routing/fallback strategy combining both is worth the added complexity.
Structured elaboration
How to evaluate. Build a held-out evaluation set of realistic queries with known-correct answers (or human-graded rubrics for open-ended ones), and measure each candidate model on: factual accuracy (does it get the answer right, and does it hallucinate on out-of-scope questions), latency (p50/p95 response time under realistic load), cost per query at your expected volume, and any user-experience metrics that matter for the product (helpfulness ratings, task completion rate in A/B tests). Benchmarks alone are not sufficient; a model can score well on a public benchmark and still perform differently on your specific domain and query distribution, so the evaluation set must reflect your actual traffic.
Vendor/integration risk. Beyond raw model quality, evaluate SLA guarantees, rate limits, data-handling and privacy terms, and how exposed you are if the vendor changes pricing, deprecates the model, or has an outage, since a third-party dependency carries operational risk that a benchmark score alone won't capture.
Combining both models. A common production pattern is routing: use the fast, cheap model for the majority of queries (the ones it handles well), and escalate to the slower, more accurate model only for queries flagged as higher-risk or lower-confidence, e.g., via the fast model's own confidence signal, query complexity heuristics, or a lightweight classifier trained on where the fast model tends to fail. This captures most of the cost savings of the cheap model while limiting factual-accuracy risk to the harder subset of queries that actually need it.
Worked example
Say the cheap model costs $0.20 per 1,000 queries and the accurate model costs $2.00 per 1,000 queries, roughly a 10x cost difference, and evaluation shows the cheap model is factually correct 92% of the time versus 98% for the expensive model on your held-out set. If a routing classifier can reliably identify the roughly 20% of queries where the cheap model is most likely to be wrong (say, questions requiring precise numeric facts or recent information) and escalate only those to the expensive model, the blended cost is 0.8×$0.20+0.2×$2.00=$0.16+$0.40=$0.56 per 1,000 queries, about a 72% cost reduction from always using the expensive model, while capturing most of its accuracy benefit specifically where it matters most.
Trade-offs & pitfalls
A routing strategy only pays off if the signal for "this query needs the accurate model" is genuinely predictive; a poorly calibrated router either escalates too much (eroding the cost savings) or too little (letting factuality-sensitive queries slip through to the cheap model). It also adds real engineering and operational complexity: two vendor integrations, two SLAs to monitor, and a routing component that itself needs to be evaluated and maintained over time. For an early-stage product without the traffic volume or engineering capacity to build and maintain a router well, a single well-chosen model, even if slightly suboptimal on cost or accuracy, is often the more pragmatic starting point, with routing revisited once volume and evaluation infrastructure justify the added complexity.
Explain the difference between a generative and a discriminative model. Give at least two concrete examples of each and describe how the choice between the two approaches affects a production system.
Sample Answer
Direct answer
A discriminative model learns the conditional distribution P(y∣x) directly: it maps an input straight to a label or decision boundary. A generative model learns (explicitly or implicitly) how the data itself was produced, either the joint P(x,y) or the marginal P(x). Because it models how data is generated, a generative model can also sample new data, something a discriminative model cannot do at all.
Structured elaboration
What each optimizes. A discriminative classifier like logistic regression or an SVM fits a boundary that separates classes as well as possible, using only the conditional likelihood of the label given the input. A generative model like a naive Bayes classifier, an autoregressive language model, a GAN, or a VAE instead tries to capture the full data-generating process, then derives P(y∣x) (if it needs to classify at all) via Bayes' rule from P(x∣y) and P(y).
Concrete examples.
- Discriminative: logistic regression for spam detection, an SVM for image classification, a discriminative fine-tuned BERT for sentiment classification.
- Generative: an autoregressive LLM producing text token by token, a GAN or diffusion model producing images, a VAE reconstructing and sampling from a learned latent space.
Three practical implications of the choice.
- Only a generative model can produce new content. If the product need is "write an email" or "draw an image," you need a generative model; a classifier cannot do this by construction.
- Discriminative models are usually more sample-efficient and accurate for pure classification. They spend their entire capacity on the decision boundary instead of also modeling the full input distribution, so with a fixed dataset size a discriminative model typically classifies better.
- Generative models carry model-misspecification risk that discriminative models mostly avoid. If your assumptions about how the data is generated are wrong (e.g., a bad choice of prior or emission distribution), a generative model's classification accuracy degrades along with its generative quality. A discriminative model that never modeled the generative process has no such assumption to get wrong.
Worked example
Take spam detection. A discriminative logistic regression sees an email's features (word counts, sender domain, links) and directly outputs P(spam∣features); it has no notion of what a "typical" spam email looks like beyond the boundary it fit. A generative naive Bayes classifier instead estimates P(word∣spam) and P(word∣not spam) for every word, then combines them with Bayes' rule to classify a new email. Because it explicitly modeled P(x∣y), that same naive Bayes model could (crudely) generate a plausible spam-sounding word distribution, while the logistic regression has no way to do so; it only ever learned where the boundary sits.
Trade-offs & pitfalls
The two are not always in tension: a generative classifier CAN outperform a discriminative one when labeled data is scarce, because it borrows statistical strength from modeling the unlabeled input distribution too. But when you have plenty of labeled data and only need a decision, the extra generative modeling is often wasted compute and can hurt calibration if the generative assumptions are even slightly wrong. The common pitfall is picking a generative model "because it's more powerful" for a task that is purely classification; if the product only ever needs a label, a well-tuned discriminative model is usually cheaper to train, faster to serve, and more accurate.
Explain adapter modules for transformer models: how they are inserted (e.g., between attention and feed-forward), how they change parameter budgets, their advantages relative to full fine-tuning, and potential drawbacks. Design a lightweight adapter architecture for sequence classification and estimate the number of extra parameters for a 1.5B parameter base model.
Sample Answer
Direct answer: Adapter modules are small bottleneck feed-forward blocks (a down-projection, a nonlinearity, and an up-projection with a residual connection) inserted into each transformer layer, most commonly right after the attention output and after the feed-forward network, so that only the adapters (and possibly the final head) need to be trained while the pretrained backbone stays frozen.
Structured elaboration:
- Insertion points: the two most common locations are between the attention block's output and its residual add-and-normalize step, and between the feed-forward network's output and its own residual add-and-normalize step, giving a typical pattern of layer output flowing through the adapter, then a residual add, before the next layer; inserting adapters inside the feed-forward network itself (between its two dense layers) is less common.
- Parameter budget: for hidden size H and adapter bottleneck dimension r, one adapter costs roughly 2Hr parameters (a down-projection of size H×r plus an up-projection of size r×H, biases are negligible); with one adapter per layer this scales as 2Hr×L across L layers, or double that with two adapters per layer (after attention and after the feed-forward network). Relative to full fine-tuning, which updates essentially 100% of the model's parameters, adapters typically add only a small fraction of extra parameters while training well under 5% of the model's total parameter count.
- Advantages relative to full fine-tuning: far cheaper storage per task (one small adapter file instead of a full checkpoint), faster training with a smaller memory footprint (far fewer gradients and optimizer states), and structurally safer against catastrophic forgetting since the backbone's weights never move, which also makes multi-task or multi-tenant setups practical by keeping one shared frozen backbone with many swappable adapters.
- Drawbacks: adapters can trail full fine-tuning's peak performance on tasks that genuinely require large representational change, they add a small but real per-layer inference cost (since, unlike LoRA, the nonlinearity between projections generally prevents merging the adapter back into the base weights), the insertion points and bottleneck size need tuning, and stacking multiple adapters (for multi-task use) can introduce interaction effects that need to be validated rather than assumed benign.
Worked example: A lightweight adapter design for sequence classification on a 1.5B-parameter base model: place one adapter after the feed-forward output of each layer (a lighter-weight choice than adapters at both attention and feed-forward positions), with a down-projection to a bottleneck dimension r, a ReLU or GeLU nonlinearity, an up-projection back to hidden size H, and a residual connection, plus a small task-specific classification head on top of the pooled output. Assuming a hidden size H≈2048 and L=24 layers (typical dimensions for a model in this parameter range): with a lightweight bottleneck of r=64 and one adapter per layer, the parameter cost is 2×2048×64=262,144 per adapter, so 262,144×24≈6.3M total, about 0.42% of the 1.5B base model. With a moderate bottleneck of r=256 and two adapters per layer, the cost per layer is 2×(2×2048×256)≈2.1M, so across 24 layers the total is roughly 50.4M, about 3.4% of the base model, both estimates verified directly by the arithmetic above.
Trade-offs and pitfalls: Choosing a bottleneck dimension that is too small can under-fit tasks needing more representational change, while choosing it unnecessarily large gives up much of the parameter-efficiency benefit without a clear performance gain, so the practical approach is to start small (for example r=64) and increase only if validation performance clearly justifies it. Unlike LoRA, most adapter designs cannot be merged back into the frozen backbone before serving because of the nonlinearity between the two projections, so their small extra inference cost is a permanent, not one-time, overhead, worth weighing against LoRA when the deployment constraint is strict per-request latency.
Explain the differences between zero-shot, one-shot, few-shot, and in-context learning in LLMs. Describe scenarios where each is preferred, and when you would reach for fine-tuning instead of relying on in-context capabilities.
Sample Answer
Direct answer
Zero-shot means asking the model to perform a task with only an instruction and no examples; one-shot gives exactly one example; few-shot gives a handful (typically 2 to a few dozen) of input-output examples in the prompt. In-context learning (ICL) is the umbrella capability that makes all three work: the model adapts its behavior to a new task purely from what is in the current prompt, with no weight updates at all. You would reach for fine-tuning instead of relying on in-context capability when the task needs to be applied at scale with tight latency/cost budgets, needs behavior more reliable than prompt-level steering can guarantee, or when the pattern is too complex or too far from the model's pre-training distribution for a handful of examples to convey.
Structured elaboration
How ICL actually works. The model was never explicitly trained to "learn from examples in a prompt" as a separate mechanism; the behavior emerges from next-token prediction at scale, where predicting the continuation of "input: X -> output: Y" patterns during pre-training exposed the model to enough structurally similar sequences that it generalizes to a novel task specified the same way at inference time, without any parameter update.
When each is preferred.
- Zero-shot is preferred when the task is common enough (or well enough described by an instruction) that the model likely saw very similar instructions during training or instruction tuning, e.g., "summarize this," "translate this to French."
- One-shot is useful mainly to pin down an output FORMAT (e.g., "here is the exact JSON shape I want") rather than to teach a genuinely new task.
- Few-shot is preferred when the task is more specific or the output format/style is unusual enough that one example is ambiguous but a handful of varied examples disambiguates it, e.g., classifying support tickets into a company-specific taxonomy.
When fine-tuning wins instead. Every example you put in a few-shot prompt costs tokens on every single request, forever, which adds real latency and dollar cost at scale; fine-tuning pays that cost once, up front, and then every inference call is short and fast. Fine-tuning is also the right call when you need the model's behavior to be reliably consistent (few-shot performance is famously sensitive to which examples you pick and what order you put them in) or when the task requires knowledge or a pattern too subtle to convey in a handful of in-prompt demonstrations.
Worked example
Suppose you're building a support-ticket triage classifier into 12 company-specific categories. Zero-shot with just category names in the instruction will likely confuse categories with overlapping language. Few-shot with 2 examples per category (24 examples) fixes most of the ambiguity, but now every classification call carries those 24 examples in the prompt, at, say, 60 tokens each, roughly 1,400 extra tokens per request purely for the examples. At 100,000 requests per day, that is 140 million extra prompt tokens per day, every day, indefinitely. Fine-tuning on a few thousand labeled tickets removes that per-request tax entirely: the fine-tuned model has "absorbed" the category boundaries into its weights, so inference goes back to a short prompt with the ticket text alone.
Trade-offs & pitfalls
The common mistake is treating fine-tuning and in-context learning as mutually exclusive; in practice teams often start with few-shot prompting to validate that the task is even learnable and to gather a labeled dataset from real usage, THEN fine-tune once volume justifies paying the one-time training cost to remove the recurring per-request example tax. Jumping straight to fine-tuning before validating the task with a cheap few-shot prototype risks spending real engineering and compute effort locking in a task definition that later turns out to be wrong or incomplete.
Design an end-to-end, privacy-compliant human-feedback and RLHF data pipeline for a production system that may collect PII or sensitive content (for example a customer-support assistant under GDPR or similar regulation). Cover redaction and annotator safety for sensitive or disallowed content, anonymization/pseudonymization and data-minimization strategies, differential privacy during training where appropriate, consent and deletion (erasure) workflows, secure labeling-platform access controls, auditability and provenance so the pipeline can be shown not to leak customer data, and the trade-off between privacy guarantees and reward-model utility.
Sample Answer
Direct answer: An end-to-end, privacy-compliant RLHF pipeline for a PII-handling production system needs PII redaction and minimization built into every stage (SFT data, preference collection, and logging), a reward model validated specifically against safety and leakage probes, a constrained RL algorithm with hard safety gates on top of the usual KL penalty, and staged rollout with automated rollback tied to safety-specific metrics, not just general quality metrics.
Structured elaboration:
- SFT data selection: source from historical support transcripts, canned responses, and policy or knowledge-base documents, run automatic PII redaction (detectors, regex, named-entity recognition) with a sampled human-verification pass on the redaction itself, remove privileged or legal content, and retain redaction metadata for later audit; keep data retention minimal, encrypted at rest, and access-controlled.
- Privacy-aware preference collection: prefer synthetic or already-redacted contexts for preference-comparison tasks; if real dialogue context is genuinely needed, apply irreversible anonymization (hashing plus surrogate tokens) and consider differential-privacy noise when aggregating labels; run the annotator environment in a compliance-appropriate sandbox with strict access controls, ephemeral tokens, and no raw data export.
- Consent and deletion (erasure) workflows: capture explicit, purpose-specific consent at the point a customer's content is collected, before it can be used in SFT data, preference labeling, or reward-model training, and version that consent record so a later audit can show what a given piece of data was and was not authorized for; build a deletion/erasure pipeline (a GDPR Article 17 requirement in this scenario) that can trace a specific user's contributions across the raw transcript store, the redacted SFT corpus, and the preference store, and that at minimum purges the raw and redacted records on request; for data that has already been folded into a trained reward model or policy checkpoint, be explicit that a single user's influence generally cannot be surgically removed from already-trained weights, so the practical mitigation is flagging affected model versions for the next scheduled retrain on the post-deletion dataset, and documenting that retraining cadence as part of the erasure SLA rather than promising instant weight-level removal.
- Reward model architecture and validation: a transformer-based scorer fine-tuned on pairwise labels via a Bradley-Terry-style loss, with uncertainty estimation (an ensemble or MC-dropout) to flag out-of-distribution inputs, validated not only on held-out preference accuracy but on SAFETY-SPECIFIC test suites, explicit PII-leakage probes and policy-violation tests, since a reward model can look well-calibrated on ordinary preferences while still failing the specific safety property this deployment needs.
- RL algorithm choice and constraints: a KL-constrained PPO variant (or DPO/a conservative-regression alternative to reduce policy-gradient instability and compute cost) with a dual-constraint structure, a soft KL penalty against the SFT baseline PLUS a hard, deterministic safety filter applied after generation that can block or rewrite an output regardless of what the reward model scored it, so a reward-model blind spot can never bypass the hard safety layer.
- Deployment and rollback: shadow-test the new policy against real traffic without serving it, then canary-roll it out gradually with automated gates on safety-specific signals (PII-detection triggers, policy-violation rate) in addition to general quality metrics, and an automated rollback to the prior stable SFT model the moment a safety threshold is exceeded, not only when overall quality drops.
- Budgeting: label-efficient formats (pairwise comparisons over absolute ratings) and active learning to prioritize which pairs to label, and cost-reducing training choices (mixed precision, gradient checkpointing, or DPO instead of full PPO) to control the RL stage's compute cost, which is typically the most expensive part of the pipeline.
Worked example: A concrete safety-gate design: even if the reward model scores a candidate response highly, a deterministic post-generation policy filter independently checks for PII patterns and disallowed content categories, and can override or rewrite the response regardless of the reward score, this dual-constraint structure (soft KL penalty for general drift plus a hard, reward-model-independent safety check) is what prevents a reward-model calibration gap specifically on rare PII-adjacent inputs from becoming a production leak.
Trade-offs and pitfalls: The dominant trade-off in a regulated PII-handling deployment (for example one subject to GDPR or a similar regulation) is conservatism versus utility, stronger filters and tighter KL constraints reduce leakage and policy-violation risk but can measurably reduce helpfulness, and the correct default here is erring conservative (accepting some utility loss) rather than the reverse, since a leaked PII incident is far costlier than a slightly less helpful response. A second pitfall is validating the reward model only on general preference accuracy without the safety-specific probes, a reward model that looks excellent on ordinary helpfulness comparisons can still have never been tested against the exact failure mode (PII leakage under an unusual phrasing) that matters most for this deployment.
Unlock Full Question Bank
Get access to all 37 Generative AI and Large Language Models interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.