Model Deployment and Inference Optimization Questions
Serving trained models efficiently in production. Covers deployment and containerization, real-time and batch serving, latency budgets, throughput and cost optimization, quantization and model compression, and online/real-time learning constraints. Emphasizes meeting production performance targets without sacrificing model quality.
Scenario: You must choose between deploying a large pretrained model (500M+ parameters) or training a smaller model plus distillation for a search re-ranking task. Business constraints: P95 latency budget 20ms, limited serving cost per QPS, and desire for highest possible relevance. Describe how you would evaluate which approach to take, including offline proxies, latency/cost estimation, and rollout strategy.
A recent model upgrade has increased tail latency (p99) from 100ms to 300ms for an image-service in production. Describe an end-to-end root-cause analysis plan spanning model, runtime, orchestration, networking, instance sizing, and request routing. Include how you'd use profiling, distributed tracing, synthetic tests, canary rollbacks, and what safety steps you'd take to minimize customer impact during diagnosis.
Describe knowledge distillation at a practical level. Explain teacher-student training setup, soft-labeling with temperature scaling, combining distillation loss with task loss, and typical scenarios where distillation improves latency or memory while preserving most accuracy. Include a short validation plan for a distilled model before production rollout.
Describe how knowledge distillation could be adapted to distill across modalities (e.g., a vision teacher to a multimodal student) or across tasks (transfer learning). What loss terms and training signals would you include to preserve cross-modal knowledge during compression?
Estimate the memory footprint and approximate floating point operations (FLOPs) for a Transformer model with ~350 million parameters, sequence length 512, and batch size 8 during inference. Clearly state assumptions (e.g., float32 weights, number of layers, hidden size) and recommend hardware choices and optimizations (quantization, model parallelism) to meet a 500ms latency SLO.
Unlock Full Question Bank
Get access to all 24 Model Deployment and Inference Optimization interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.