Deep Learning: Neural Networks and Architectures Questions
How deep neural networks are built and trained. Covers network fundamentals, activation functions and non-linearity, loss function selection, backpropagation, optimization and learning-rate behavior, and diagnosing vanishing or exploding gradients, along with the major architecture families: convolutional networks for images and recurrent networks, LSTMs, and GRUs for sequences. Emphasizes both how to train deep networks stably and how to match an architecture family to the structure of the data.
Describe how embedding layers work for categorical variables in a neural network: how to choose embedding dimensionality, handle unseen categories at inference, and integrate embeddings with numerical features in a feedforward model.
Sample Answer
Direct answer
An embedding layer is just a lookup table, a learned matrix of shape (number of categories, embedding dimension), that converts a discrete category ID into a dense vector the rest of the network can process like any other numeric feature.
Structured elaboration
Representing inputs: map each category to an integer ID in [0,N), reserving one dedicated ID for unknown or missing categories; the embedding layer itself is an N×d matrix, and looking up an ID simply selects that row.
Choosing dimensionality d: a common rule of thumb is d≈min(50,⌈6N1/4⌉), growing slowly with the number of distinct categories N; in practice, start small (8 to 64) and tune against validation performance, since an oversized embedding for a low-cardinality feature mostly adds overfitting risk and wasted parameters without adding real capacity.
Handling unseen categories at inference: reserve a dedicated "unknown" ID during training (routing sufficiently rare categories to it during training itself, not only at inference, so its embedding is actually learned from real examples rather than starting cold at serving time), or use feature hashing (mapping category values into a fixed number of hash buckets), which guarantees every possible future category maps to SOME existing embedding, at the cost of occasional hash collisions between genuinely different categories.
Integrating with numerical features: normalize numerical features first (their raw scale is generally very different from an embedding's learned scale), concatenate all embedding vectors together with the normalized numeric features into one combined vector, and feed that into the feedforward network's remaining dense layers.
Worked example
import torch
embs = [emb_layer(ids[:, i]) for i, emb_layer in enumerate(emb_layers)]
x_cat = torch.cat(embs, dim=1) # (batch, sum of embedding dims)
x_num = num_bn(num_tensor) # normalized numeric features
x = torch.cat([x_cat, x_num], dim=1)
out = mlp(x)
For a categorical feature with N=1000 distinct values, the rule of thumb gives d≈min(50,6×10000.25)=min(50,6×5.62)=min(50,33.7)≈34, a moderate embedding size that is neither a single scalar (too little capacity to capture 1000 distinct categories' relationships) nor a needlessly huge vector for a feature this size.
Trade-offs & pitfalls
A common mistake is treating the "unknown" bucket as a purely inference-time fallback, never actually exposing the model to it during training; if the model has never seen the unknown-ID embedding receive a real gradient update, it starts from an untrained, effectively random initialization at exactly the moment (an unseen category at inference) when a well-calibrated fallback matters most. A second common gap is skipping normalization of the accompanying numeric features before concatenation; embeddings are typically initialized and trained to occupy a specific, learned numeric range, and unnormalized numeric features with a very different natural scale can dominate or be dominated by the embedding features purely due to scale, not genuine predictive importance.
Explain the architecture of an LSTM cell: input gate, forget gate, output gate, candidate cell, and cell state. Give the forward-pass equations and explain why the gating structure preserves long-range dependencies compared to a vanilla RNN.
Sample Answer
Direct answer
An LSTM cell adds three gates and a separate cell state on top of a plain recurrent unit, letting the network learn WHEN to remember, overwrite, or expose information, which is what lets gradients (and information) survive across many more timesteps than a vanilla RNN can manage.
Structured elaboration
Given input xt, previous hidden state ht−1, and previous cell state ct−1:
ft=σ(Wf[ht−1,xt]+bf)(forget gate: how much of ct−1 to keep)
it=σ(Wi[ht−1,xt]+bi)(input gate: how much new content to write)
c~t=tanh(Wc[ht−1,xt]+bc)(candidate cell content)
ct=ft⊙ct−1+it⊙c~t(cell state update)
ot=σ(Wo[ht−1,xt]+bo)(output gate: how much of ct to expose)
ht=ot⊙tanh(ct)
Sigmoid gates produce values in (0,1), acting as soft, learnable on/off switches; tanh keeps candidate content and the exposed hidden state in (−1,1), centered around zero, which helps gradient-based optimization.
Why the gating structure preserves long-range dependencies: the cell-state update is ADDITIVE (ft⊙ct−1+it⊙c~t), not a repeated multiplicative transform of the kind that causes vanishing gradients in a vanilla RNN's hidden-state recurrence. When the forget gate is close to 1, the gradient with respect to ct−1 passes through nearly unattenuated (multiplied by approximately 1, not by a small, saturating activation derivative), so error signal can flow back across many timesteps largely intact, as long as the network has learned to keep the relevant forget gates open along the way.
Worked example
Suppose at some timestep ft=0.95 (mostly remember), it=0.1 (write a little new content), ct−1=2.0, c~t=0.5: ct=0.95(2.0)+0.1(0.5)=1.9+0.05=1.95, close to the previous cell value, as expected when the forget gate is near 1. Contrast with ft=0.1 (mostly forget), same other values: ct=0.1(2.0)+0.1(0.5)=0.2+0.05=0.25, showing the cell state genuinely discarding most of its prior content when the network has learned that forgetting is appropriate at this step.
Trade-offs & pitfalls
A common oversimplification is describing the cell state as "immune" to vanishing gradients; it is much MORE resistant than a vanilla RNN's hidden state, but the forget gate itself is still a learned, saturating sigmoid, so if training pushes forget gates toward the extremes (stuck near 0) the additive benefit is undermined. A practical mitigation often used is initializing the forget-gate bias to a small positive value (e.g. +1), biasing the network toward remembering by default early in training, which tends to make long-range dependencies easier to learn from the start rather than relying on training to discover that remembering is useful.
Compare SGD, SGD with momentum, RMSProp, Adam, and AdamW: the high-level update rule, sensitivity to learning rate, convergence behavior, and generalization tendencies. Explain why AdamW separates weight decay from the adaptive update and when you would recommend each optimizer.
Sample Answer
Direct answer
SGD, SGD with momentum, RMSProp, Adam, and AdamW form a spectrum from simple-but-slow to adaptive-but-generalization-risky; AdamW is usually the safe modern default, but SGD with momentum still wins on final generalization for many vision tasks when you have the compute budget to tune it properly.
Structured elaboration
| Optimizer | Update rule (high level) | Learning-rate sensitivity | Convergence | Generalization tendency |
|---|---|---|---|---|
| SGD | θ←θ−η∇θL | High, needs careful scheduling | Slow, can oscillate in narrow ravines | Often the best final generalization if properly tuned |
| SGD + momentum | v←μv+η∇θL; θ←θ−v | Lower than plain SGD | Faster, dampens oscillation | Comparable to plain SGD, usually strong |
| RMSProp | s←ρs+(1−ρ)∇2; θ←θ−η∇/(s+ϵ) | Low, per-parameter adaptive scaling | Fast early progress; can bounce near minima | Mixed; good for noisy/sparse gradients, sometimes weaker than SGD on vision |
| Adam | Adds bias-corrected first-moment estimate m^ on top of RMSProp's second moment v^: θ←θ−ηm^/(v^+ϵ) | Low, robust default η≈10−3 | Very fast initial convergence | Can generalize worse than SGD on some vision tasks; tends toward sharper minima |
| AdamW | Same as Adam, but with weight decay applied as a separate, decoupled term: θ←θ(1−ηλ)−ηm^/(v^+ϵ) | Similar to Adam | Similar to Adam | Improves on Adam's generalization by restoring a consistent weight-decay effect |
Why AdamW decouples weight decay: naive L2 regularization adds λw into the gradient before the adaptive rescaling by v^, so a parameter with a large accumulated second moment gets its decay term shrunk by that same adaptive factor, which means differently-scaled parameters are decayed by inconsistent effective amounts, undermining the regularization's intent. AdamW instead applies the decay directly to the parameter, outside the adaptive-rescaling step, so every parameter is shrunk by the same proportional amount regardless of its gradient history.
When to recommend each: SGD with momentum for large-scale vision training and any setting where final generalization matters more than wall-clock convergence speed, paired with a learning-rate schedule and enough compute budget to tune it. Adam or AdamW for transformers, NLP, and rapid prototyping, where fast, robust initial convergence with less tuning effort matters more, and where AdamW is close to a strict improvement over Adam at essentially no extra cost. RMSProp and plain AdaGrad are mostly of historical interest now (RMSProp still shows up for RNNs and some reinforcement-learning setups); AdaGrad's monotonically-shrinking accumulated second moment is its own known shortcoming since it eventually stalls the effective learning rate to near zero, which is exactly the failure mode Adam's exponential moving average was designed to fix. For very large batch, large-scale distributed training, plain SGD with momentum plus a linear learning-rate scaling rule and warmup remains a common, well-understood choice, in part because its update rule is simpler to reason about and synchronize across many workers.
Worked example
One Adam update step with β1=0.9, β2=0.999, η=10−3, ϵ=10−8, starting from m0=0, v0=0, at step t=1 with gradient g=0.5:
m1=0.9(0)+0.1(0.5)=0.05, v1=0.999(0)+0.001(0.25)=0.00025
Bias-corrected: m^1=0.05/(1−0.91)=0.5, v^1=0.00025/(1−0.9991)=0.25
Update: θ←θ−10−3×0.5/(0.25+10−8)=θ−10−3×1.0=θ−0.001
The bias correction matters here precisely because at t=1 the raw moments are heavily biased toward zero (their initialization); without it, the first update would be far smaller than intended.
Trade-offs & pitfalls
A common mistake is treating "switch to Adam" as a strict upgrade in every setting; on several large vision benchmarks, SGD with momentum plus a well-tuned schedule still reaches better final validation accuracy than Adam, at the cost of more tuning effort. A second pitfall is applying naive L2 weight decay under Adam and expecting it to behave like classical weight decay under SGD; because of the adaptive-rescaling interaction described above, it will not, which is the entire motivation for AdamW.
List and explain the common regularization techniques used in deep learning: dropout, weight decay (L1/L2), data augmentation (including mixup/cutmix), early stopping, batch normalization as an implicit regularizer, and label smoothing. For each, describe the mechanism and a rule of thumb for when to apply it.
Sample Answer
Direct answer
Dropout, weight decay, data augmentation, early stopping, batch normalization, and label smoothing all reduce overfitting, but each does so by constraining a different part of the system: activations, weight magnitudes, the input distribution, training duration, internal statistics, or target confidence.
Structured elaboration
| Technique | Mechanism | Rule of thumb |
|---|---|---|
| Dropout | Randomly zeroes a fraction of activations each forward pass, preventing units from co-adapting and approximating an implicit ensemble of thinned subnetworks | Apply in fully-connected layers; use sparingly or not at all alongside heavy batch normalization, which already has a regularizing effect |
| Weight decay (L1/L2) | Adds a penalty on weight magnitude to the loss (λ∑w2 for L2, λ∑∣w∣ for L1), discouraging large weights and favoring smoother learned functions; L1 additionally drives many weights exactly to zero (sparsity), L2 shrinks them all proportionally | Use L2/weight decay as a near-default; reach for L1 specifically when you want automatic feature selection or a sparse model |
| Data augmentation | Expands the effective training distribution with label-preserving transformations (crop, flip, color jitter, mixup, cutmix) | Use domain-appropriate transforms; this is often the single highest-leverage regularizer in vision and audio |
| Early stopping | Halts training once validation performance stops improving, effectively limiting how long the model can keep fitting training-set idiosyncrasies | Apply almost always; cheap, robust, needs only a validation set and a patience parameter |
| Batch normalization (implicit) | Normalizes each layer's inputs using per-batch statistics, which both stabilizes training and introduces a small amount of noise (since batch statistics vary batch to batch) that acts as a mild regularizer | Use for training stability first; treat any regularizing side-effect as a bonus, not the primary reason to add it |
| Label smoothing | Replaces one-hot targets with a softened distribution, discouraging the network from driving logits to extreme, overconfident values | Use when calibration matters or the label set is large; avoid when the task genuinely needs near-certain, sharply confident predictions |
Worked example
L2 regularization's effect on the loss, made concrete: for a loss L(θ)=L0(θ)+λ∥θ∥2, the gradient becomes ∇L=∇L0+2λθ, so an SGD update becomes θ←θ−η∇L0−2ηλθ=(1−2ηλ)θ−η∇L0. The (1−2ηλ) factor shrinks every weight multiplicatively toward zero on every step, independent of the data-driven gradient, which is exactly the "decay" in weight decay. For η=0.01, λ=0.01: each step shrinks weights by a factor of 1−2(0.01)(0.01)=0.9998, a small but compounding effect over thousands of steps.
Trade-offs & pitfalls
A common mistake is stacking many of these techniques at maximum strength simultaneously and then being unable to tell which one is responsible for a given change in validation performance; ablate one at a time when tuning. A second pitfall specific to L1/L2 under adaptive optimizers like Adam is that naive L2 does not behave like true weight decay once gradients are being adaptively rescaled per-parameter, which is exactly the motivation for AdamW's decoupled weight decay.
Describe strategies for training when labeled data is scarce: self-supervised pretraining, semi-supervised learning, data augmentation, and transfer learning. How would you use learning curves and active learning to decide if you have enough labeled data, and how would you validate gains without confirmation bias?
Sample Answer
Direct answer
When labeled data is scarce, self-supervised pretraining, semi-supervised learning, augmentation, and transfer learning all extract extra signal from the SAME small labeled set (or from unlabeled data you already have), but validating that any of them genuinely helped requires a validation protocol specifically designed to resist confirmation bias, since pseudo-labeling in particular can easily produce numbers that look like progress but aren't.
Structured elaboration
Self-supervised pretraining: learn general representations from UNLABELED data first (contrastive methods like SimCLR/MoCo, or masked-reconstruction methods like MAE), giving the network a strong starting point before it ever sees the scarce labels; a quick way to sanity-check whether this pretraining actually captured something useful is LINEAR PROBING (freezing the pretrained representation and training only a simple linear classifier on top), which is cheap and gives an early signal before committing to full fine-tuning.
Transfer learning: fine-tune (or use as a frozen feature extractor) a checkpoint already pretrained on a large, related dataset; compare a frozen-feature linear probe against full fine-tuning specifically to gauge how much genuine domain shift exists between the pretraining data and your task.
Semi-supervised learning: consistency regularization (enforcing that predictions stay stable under input augmentation, as in Mean Teacher or FixMatch) and pseudo-labeling (assigning high-confidence labels to unlabeled examples and training on them) both use UNLABELED data beyond the scarce labeled set, typically via a teacher-student or exponential-moving-average setup that smooths the targets rather than trusting the model's own raw, possibly-overconfident predictions directly.
Data augmentation: label-preserving transformations (and stronger policies like RandAugment or MixUp/CutMix) synthetically expand the effective diversity of the scarce labeled set.
Using learning curves and active learning to decide if you have enough data: plot validation performance against labeled-dataset SIZE (training repeatedly on increasing subsets); a curve that is still climbing steeply at your current dataset size is direct evidence more labels would help meaningfully, while a curve that has visibly flattened suggests you're near the point of diminishing returns FOR THIS architecture and task. Active learning (having the model flag its most UNCERTAIN unlabeled examples for human labeling next) directs limited labeling budget toward the examples likely to teach the model the most, rather than labeling more data uniformly at random.
Worked example
A concrete anti-confirmation-bias validation protocol for pseudo-labeling specifically: reserve a CLEAN labeled validation set that NEVER participates in generating pseudo-labels and is never touched by the semi-supervised training loop at all; run an explicit ABLATION comparing (a) labels-only baseline, (b) labels plus self-supervised pretraining, and (c) labels plus pretraining plus pseudo-labeling, all evaluated on that SAME untouched clean validation set; and require the reported improvement to be consistent across multiple random seeds (not a single lucky run) before accepting it as real, ideally with a statistical test (a paired comparison or bootstrap confidence interval) rather than a single point-estimate comparison.
Trade-offs & pitfalls
The single most common confirmation-bias trap is evaluating a pseudo-labeling pipeline on a validation set that WAS INVOLVED, even indirectly, in generating the pseudo-labels being trained on; this systematically inflates the reported gain, since the model has effectively been validated against data whose "ground truth" the model itself helped produce. A second common mistake is accepting a pseudo-labeling gain from a SINGLE training run without checking it holds up across seeds or via a proper statistical comparison; semi-supervised methods can show meaningfully higher run-to-run variance than fully-supervised training, and a single favorable run is weak evidence on its own.
Unlock Full Question Bank
Get access to all Deep Learning: Neural Networks and Architectures interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.