Model Training Infrastructure and Distributed Training Questions

Scaling model training across hardware and time. Covers GPU/accelerator considerations, data and model parallelism, distributed and large-scale training, experiment tracking and training infrastructure, and the training-versus-inference compute tradeoff. Focuses on the systems and resource decisions that make large-model training feasible.

MediumTechnical
81 practiced

Design an experiment-tracking schema and REST API to store experiment runs. Provide a sample JSON schema or SQL table design that captures run_id, start/end timestamps, git_hash, dataset_version, hyperparameters (nested), metrics per epoch, artifact URIs, and tags. Explain indexing and query patterns for retrieving top-k runs by metric and for filtering by hyperparameter ranges.

MediumTechnical
99 practiced

Technical/theoretical (medium): Explain how Differential Privacy via DP-SGD would be applied in large-scale distributed training. List practical engineering challenges (noise scale, clipping, compiler/runtime support) and mitigation strategies (microbatching, accounting, custom kernels).

MediumTechnical
104 practiced

Explain DeepSpeed ZeRO's stages 1–3 (optimizer-state sharding, gradient sharding, parameter sharding). For each stage, describe the memory savings achieved, the additional communication patterns or overhead introduced, and recommended scenarios (model sizes and GPU counts) where each stage is most beneficial.

MediumTechnical
104 practiced

Write Python pseudocode for a mini-batch training loop using PyTorch that supports checkpointing, early stopping based on validation loss, and resuming from a saved checkpoint. Focus on structure: saving state_dicts, optimizer state, epoch counter, and logic for resume and early stop.

MediumTechnical
96 practiced

Compare single-node training versus distributed data-parallel training from the perspective of numerical reproducibility. What causes non-determinism across runs (e.g., reduction ordering, non-associativity of FP ops), and what engineering techniques can you apply to increase reproducibility in production?

Unlock Full Question Bank

Get access to all Model Training Infrastructure and Distributed Training interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.