LLM Fine-Tuning and Alignment Questions

Adapting foundation models to specific tasks and desired behavior. Covers transfer learning and using pretrained models, full and parameter-efficient fine-tuning, instruction tuning, and alignment methods such as RLHF and preference optimization. Focuses on when and how to customize a base model rather than prompt it, and the data and compute tradeoffs involved.

HardSystem Design
54 practiced

Architect an end-to-end RLHF training platform or pipeline for a production instruction-following assistant at scale (for example 100M preference pairs, supporting daily fine-tuning runs). Describe the major components (data ingestion, annotation service, preference store, reward-model training, policy-optimization cluster, artifact repository, serving layer, monitoring), data flow, sharding/partitioning strategies, and main compute/storage considerations and cost-saving opportunities (GPU/TPU sizing, checkpoint frequency and retention, throughput needs for offline and online scoring).

HardTechnical
64 practiced

Propose engineering and evaluation strategies to detect and mitigate adversarial inputs that exploit a reward model or policy post-RLHF, for example prompt injection, token stuffing, or paraphrase attacks designed to game the reward or trigger a reward increase. What would you build to catch these attacks before and after deployment, and how would you validate that your defenses actually work rather than just seeming to?

EasyTechnical
64 practiced

In Python, implement a function that converts a list of pairwise preference records into training pairs for a reward model. Input: list of tuples (prompt, completion_a, completion_b, preferred) where preferred is 'A' or 'B'. Output: list of examples [(input_text, label)] where label is 1 if A preferred else 0, and input_text encodes prompt and both completions using the format: 'PROMPT: <prompt>
A: <completion_a>
B: <completion_b>'. Document assumptions.

HardTechnical
56 practiced

Discuss how RLHF can be combined with offline reinforcement learning techniques to leverage large historical logs for alignment while avoiding live rollouts. Describe algorithmic considerations and how to validate offline-learned policies before limited online exposure.

HardTechnical
69 practiced

You see an increase in hallucinations after fine-tuning an LLM on domain-specific QA. Propose a systematic debugging and mitigation plan: experiments to isolate the cause, dataset checks, training interventions, and runtime techniques to reduce hallucinations.

Unlock Full Question Bank

Get access to all LLM Fine-Tuning and Alignment interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.