Model Deployment and Inference Optimization Questions

Serving trained models efficiently in production. Covers deployment and containerization, real-time and batch serving, latency budgets, throughput and cost optimization, quantization and model compression, and online/real-time learning constraints. Emphasizes meeting production performance targets without sacrificing model quality.

HardTechnical
24 practiced

Explain how compiler toolchains like XLA, TVM, and Glow can be used to optimize models for specific hardware. Compare the kinds of optimizations they perform (operator fusion, layout transformation, autotuning) and practical considerations for using them in a production model optimization pipeline.

HardTechnical
19 practiced

You're serving a vision model that depends on a custom CUDA kernel not supported by ONNX Runtime. Describe how you'd integrate that custom kernel into a production inference pipeline: building and packaging the kernel, runtime registration and versioning, CI tests, cross-platform builds, and fallback to CPU or alternate kernels if GPU support is unavailable.

HardTechnical
22 practiced

Implement a Python function that searches for optimal per-layer bitwidths for quantization using a simple greedy heuristic: iterate layers, try lowering bitwidth (e.g., 8 → 4) if validation metric stays within threshold, and lock choices. Provide interface, pseudocode, and complexity analysis for a model with L layers.

MediumTechnical
19 practiced

Tail latency (p99) for generation is unacceptable. Discuss specific techniques to reduce p99 for autoregressive generation: speculative decoding, early-exit/prediction confidence, hedged requests, prioritized scheduling, and caching. For each technique, explain how it impacts cost, average latency, and worst-case latency.

HardSystem Design
22 practiced

Design an inference architecture to serve a 100B-parameter LLM for real-time chat with a 50ms target latency per token. Include choices for model sharding (tensor vs pipeline), inter-GPU networking (NVLink/NCCL), activation checkpointing/offloading, memory and compute balancing, batching strategies, and how caching of past key/value tensors is handled for chat sessions. Discuss hardware and cost trade-offs.

Unlock Full Question Bank

Get access to all 11 Model Deployment and Inference Optimization interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.