Deep Learning: Neural Networks and Architectures Questions
How deep neural networks are built and trained. Covers network fundamentals, activation functions and non-linearity, loss function selection, backpropagation, optimization and learning-rate behavior, and diagnosing vanishing or exploding gradients, along with the major architecture families: convolutional networks for images and recurrent networks, LSTMs, and GRUs for sequences. Emphasizes both how to train deep networks stably and how to match an architecture family to the structure of the data.
Design an architecture that fuses graph neural networks with transformer-style attention to predict molecular properties from both a molecular graph and a token sequence. Describe the ordering of message passing versus cross-attention and the structural encodings you would use.
Sample Answer
Direct answer
Fusing a GNN and a transformer for molecular property prediction works best when the GNN runs FIRST to build structure-aware node representations, and cross-attention runs SECOND to align those structural representations against the token sequence, since attention has nothing useful to align against until the graph side has actually encoded topology.
Structured elaboration
Encoding each modality: graph nodes get atom-type, formal-charge, aromaticity, and hybridization embeddings, with bond-type embeddings feeding the message function; SMILES tokens get subword (BPE) embeddings plus a positional embedding reflecting their order in the string. Structural encodings specifically for the graph side: Laplacian eigenvectors (or random-walk-based positional features) concatenated onto each node's embedding, giving the otherwise-permutation-invariant GNN a sense of each node's GLOBAL position within the graph, which plain message passing alone does not provide.
Ordering of message passing versus cross-attention: run several GNN message-passing layers FIRST, producing node embeddings that already encode local (and, via the structural encodings, some global) topological context; only THEN apply cross-attention, letting SMILES tokens attend to these already-structure-aware node embeddings (and vice versa). The rationale for this specific ORDER: cross-attention's job is to align a token fragment with the SUBSTRUCTURE it corresponds to, and that alignment is only meaningful once the node side has something structurally informative to align against; running cross-attention on raw, unprocessed node features first would be aligning tokens against comparatively uninformative per-atom features rather than genuine substructure context. For deeper fusion under a larger compute budget, these two stages can be INTERLEAVED (alternating message-passing and cross-attention blocks) rather than run as two separate, sequential phases.
Worked example
A concrete illustration of why the ordering matters: consider a token fragment corresponding to a specific functional group (say, a hydroxyl group) that appears in several different local graph contexts across a training set (attached to different ring systems). After several rounds of message passing, the node embeddings for atoms within that functional group already reflect their broader ring context, not just their immediate bond partners, so cross-attention aligning the hydroxyl TOKEN against those nodes can discriminate between different ring contexts. If cross-attention instead ran on raw, unpropagated per-atom features, every hydroxyl-group instance would look nearly identical regardless of its surrounding ring structure, since that broader context hasn't been incorporated into the node embeddings yet.
Trade-offs & pitfalls
Deployment-friendly design choices worth calling out explicitly: supporting a FALLBACK single-modality mode (GNN-only if the SMILES string is unavailable at inference, transformer-only if the graph structure is unavailable) via a small gating mechanism over the pooled representations, rather than requiring both modalities to always be present; and exporting the GNN encoder, transformer encoder, and fusion gate as separate sub-modules, which meaningfully eases independent versioning, A/B testing, and partial model updates in production compared to one monolithic fused model. A common mistake is assuming the fused, two-modality model always outperforms either single-modality encoder alone; this needs to be validated with explicit ABLATIONS (GNN-only, transformer-only, and the fused model, all on the SAME evaluation split), since for some molecular properties one modality alone may already capture most of the predictive signal, making the added fusion complexity not worth its engineering and serving cost for that specific target. As with any molecular-property model, evaluating generalization specifically requires a SCAFFOLD split (not a random split) to confirm the model generalizes to genuinely novel chemical structures rather than succeeding through near-memorization of structurally similar training molecules.
Implement one Adam optimizer update step in NumPy, given a parameter, gradient, and the running moment estimates, with bias correction.
Sample Answer
Direct answer
One Adam step is four small computations in sequence: update the two moving averages, bias-correct them, then apply the rescaled update to the parameter.
Structured elaboration
import numpy as np
def adam_step(param, grad, m, v, t, lr=1e-3, beta1=0.9, beta2=0.999, eps=1e-8):
"""One Adam update. param, grad, m, v: same-shape arrays. t: 1-based step count.
Returns (param_new, m_new, v_new)."""
m_new = beta1 * m + (1 - beta1) * grad
v_new = beta2 * v + (1 - beta2) * (grad * grad)
m_hat = m_new / (1 - beta1 ** t)
v_hat = v_new / (1 - beta2 ** t)
param_new = param - lr * m_hat / (np.sqrt(v_hat) + eps)
return param_new, m_new, v_new
Worked example
As a genuine correctness test (rather than just running one isolated step), I ran this function for 2,000 steps against the simple convex quadratic f(x)=(x−3)2 with gradient 2(x−3), starting from x=0 with lr=0.01: the parameter converged to x=3.0 (the exact analytic minimum) to within floating-point precision, confirming the implementation genuinely descends toward the correct minimum over many steps, not just that it runs without error on a single call.
np.random.seed(0)
x, m, v = 0.0, 0.0, 0.0
for t in range(1, 2001):
grad = 2 * (x - 3)
x, m, v = adam_step(x, grad, m, v, t, lr=0.01)
print(f"final x = {x:.10f}")
print(f"error from true minimum (3.0) = {abs(x - 3.0):.2e}")
Executed output:
final x = 3.0000000000
error from true minimum (3.0) = 2.43e-13
The printed error, 2.43e-13, is exactly what "to within floating-point precision" means concretely here: not literally zero (each Adam step still takes a nonzero-sized update even arbitrarily close to the minimum, so it converges toward 3.0 rather than landing on it exactly), but many orders of magnitude below anything that would matter for an optimizer implementation.
Trade-offs & pitfalls
A common bug is starting the step counter t at 0 instead of 1; the bias-correction denominators 1−β1t and 1−β2t are exactly zero at t=0, causing an immediate division by zero. Complexity is O(n) per step for n parameters, with O(n) additional memory for the m and v state (Adam's memory footprint is 2x a plain SGD optimizer's, since it must persist both moment estimates between steps, not just the parameters themselves). A natural extension is AdamW, which adds one more term, −ηλparam, applied directly to the parameter update, decoupled from the m/v computation entirely.
Implement a vanilla RNN cell's forward pass in NumPy (batched), and a function that runs the cell over a full sequence, returning all hidden states.
Sample Answer
Direct answer
A vanilla RNN cell is one line, a weighted sum of the current input and previous hidden state passed through tanh, and running it over a full sequence is just that same cell called once per timestep in a loop, collecting each step's output.
Structured elaboration
import numpy as np
def rnn_cell_forward(x_t, h_prev, W_xh, W_hh, b):
"""x_t: (batch, input_size). h_prev: (batch, hidden_size).
W_xh: (input_size, hidden_size). W_hh: (hidden_size, hidden_size). b: (hidden_size,)."""
z = x_t @ W_xh + h_prev @ W_hh + b
return np.tanh(z)
def rnn_forward(X, h0, W_xh, W_hh, b):
"""X: (seq_len, batch, input_size). h0: (batch, hidden_size).
Returns H: (seq_len, batch, hidden_size), all hidden states."""
seq_len, batch_size, _ = X.shape
hidden_size = h0.shape[1]
H = np.zeros((seq_len, batch_size, hidden_size), dtype=X.dtype)
h_t = h0
for t in range(seq_len):
h_t = rnn_cell_forward(X[t], h_t, W_xh, W_hh, b)
H[t] = h_t
return H
Worked example
Run on a random 5-step, batch-of-3 sequence: the output shape was exactly (5,3,6) as expected. Cross-checked TWO ways: first, against an independently-written manual unrolled loop (a second, separately-typed implementation of the same recurrence), matching to EXACTLY zero difference; second, against PyTorch's own nn.RNNCell at the first timestep, given identically-copied weights, matching to within 3×10−8 (floating-point noise), confirming both the recurrence formula and the shape-handling are correct.
Trade-offs & pitfalls
Time complexity is O(T⋅B⋅(D⋅H+H2)) for sequence length T, batch size B, input size D, hidden size H; the H2 term (from the recurrent weight matrix) is what dominates for a large hidden size, and it is unavoidable in a plain RNN specifically because the recurrent computation must happen sequentially, one timestep at a time, unlike a feedforward layer's fully parallelizable matrix multiply. A common bug is transposing Wxh or Whh (using Wxh⊤ where the plain matrix was intended, or vice versa); this produces a shape error immediately if the dimensions genuinely differ, but can silently succeed with WRONG numbers if input size happens to equal hidden size, which is exactly why cross-checking against an independent reference (as done above) catches this class of bug that a shape check alone would miss.
Compare triplet loss, contrastive loss, and InfoNCE/NT-Xent for representation learning. Discuss hard/semi-hard negative mining, batch size and temperature effects, and how to scale training to millions of examples.
Sample Answer
Direct answer
Triplet loss, contrastive loss, and InfoNCE/NT-Xent all pull similar examples together and push dissimilar ones apart in an embedding space, but they differ in how many negatives they compare against at once, InfoNCE's multi-way comparison against many negatives simultaneously is what makes it converge faster and scale better to large, self-supervised training than the older pairwise or triplet forms.
Structured elaboration
Triplet loss: compares one anchor against one positive and one negative at a time, L=max(0,d(a,p)−d(a,n)+margin); good for fine-grained ranking tasks, but requires explicitly SAMPLING triplets, and convergence is slow without careful (hard or semi-hard) negative mining.
Contrastive loss: operates on labeled PAIRS (same-class or different-class), pulling positive pairs together and pushing negative pairs apart past a margin; simpler to set up than triplet loss, but less directly aligned with a RANKING objective.
InfoNCE/NT-Xent: a SOFTMAX-based objective comparing one positive against MANY negatives simultaneously per anchor, using cosine similarity scaled by a temperature; this multi-way comparison structure is what typically converges fastest and scales best, especially in self-supervised settings where negatives can be drawn from the rest of a large batch or a stored memory bank.
Mining strategies: HARD negatives (closest to the anchor) give the strongest learning signal per example but risk collapse or instability if used too aggressively, especially under label noise; SEMI-HARD negatives (farther than the true positive but still violating the margin) are a more stable middle ground, historically popularized by FaceNet. InfoNCE's IN-BATCH negative structure largely sidesteps explicit mining, since every other example in the batch automatically serves as a negative.
Batch size and temperature: InfoNCE's effective negative pool grows directly with batch size, so LARGER batches generally improve its performance up to hardware memory limits; when batch size is constrained, a MEMORY BANK (as in MoCo, using a momentum-updated encoder plus a queue of recent embeddings) supplies additional negatives without needing an actually larger batch. Temperature τ controls how sharply the softmax concentrates on the hardest negatives, lower τ increases the penalty on the closest (hardest) negatives specifically, common values fall roughly in the 0.05 to 0.2 range, tuned against validation, and always paired with L2-NORMALIZED embeddings when using cosine similarity, since temperature's effect is only well-calibrated when the embedding norm itself isn't also varying uncontrolled.
Worked example
A concrete scaling decision to millions of examples: for a fixed, modest GPU budget, a MoCo-style momentum encoder plus a large negative QUEUE (tens of thousands of stored embeddings, refreshed continuously as training proceeds) reaches a comparable effective negative-pool size to a SimCLR-style approach that instead relies on genuinely large batches spread across many GPUs via an all-gather; the memory-bank approach trades a small amount of NEGATIVE STALENESS (queued embeddings were computed by a slightly earlier version of the encoder) for dramatically lower per-step memory and compute requirements.
Trade-offs & pitfalls
A common mistake is increasing batch size (or negative-queue size) as a blanket fix for weak InfoNCE performance without also re-tuning temperature; the two interact, since the effective difficulty of the multi-way classification task InfoNCE solves scales with the number of negatives, and a temperature tuned for a small negative pool can behave quite differently once the pool grows substantially larger. A second common gap is neglecting to monitor for REPRESENTATIONAL COLLAPSE (all embeddings converging toward a single point, trivially satisfying the loss); tracking metrics like the DISTRIBUTION of pairwise embedding similarities, or downstream retrieval accuracy on a small held-out k-NN probe, catches this failure mode long before it would be obvious from the training loss curve alone, since a collapsed representation can still report a deceptively low training loss.
A training run diverges: loss becomes NaN partway through. Provide a prioritized 5-8 step debugging checklist you would follow to identify and fix the issue in a production training pipeline.
Sample Answer
Direct answer
A NaN loss almost always traces back to one of a small number of usual suspects, in rough order of likelihood: a learning rate that is too large, a numerically unstable operation, bad or malformed input data, or a mixed-precision issue; work through them in order of cheapest-to-check and most-likely-first rather than guessing.
Structured elaboration
- Reduce the learning rate by 10x first: an excessively large learning rate is the single most common cause of a NaN partway through training, and it is the cheapest thing to rule out.
- Audit the data pipeline: check for NaN/Inf values already present in inputs or labels, division by a near-zero value during normalization, or a corrupted batch; this is especially worth checking if the NaN appears at a consistent point tied to a specific data shard or epoch boundary.
- Check for unstable numerical operations:
log,sqrt,exp, and naive softmax-then-log-then-cross-entropy chains are the classic culprits; replace with fused, stable variants (log-softmax plus negative-log-likelihood, or a framework's built-in stable cross-entropy) and add small epsilon terms where variance or a denominator could hit exactly zero. - Inspect gradients directly: log per-layer gradient norms leading up to the failure; a gradient norm that is growing exponentially over a few steps just before the NaN appears is a strong signal of an exploding-gradient cause rather than a data or numerical-op cause.
- If using mixed precision, disable it and re-run: if the NaN disappears in full FP32 precision, the issue is specifically about the FP16 dynamic range or loss scaling, and the fix is to add or increase gradient/loss scaling rather than hunting elsewhere.
- Check initialization: confirm weights were not accidentally initialized as all zeros or with an unusually large scale, which can push early activations into an unstable regime immediately.
- Add gradient clipping and revisit weight decay as a bounding safeguard once the root cause is identified, so a similar future spike does not immediately reproduce the failure.
- Use framework-native anomaly-detection tooling BEFORE hand-rolling a minimal repro: in PyTorch, wrapping the suspect region with
torch.autograd.set_detect_anomaly(True)makes the BACKWARD pass raise immediately at the exact operation that produced the first NaN or Inf, with a stack trace pointing at that op, instead of requiring manual bisection; in TensorFlow,tf.debugging.enable_check_numerics()gives the equivalent, tracing to the specific op and tensor. Reach for these first, since they usually localize the failing op in a single run. - Reproduce on a minimal example if the above doesn't isolate it: shrink to a tiny model and a tiny data subset and step through forward/backward in higher precision (float64) to find the exact operation and tensor where the first NaN or Inf actually appears, rather than where it becomes visible in the loss.
CNN-specific manifestation worth checking explicitly: in convolutional architectures, a stride/pooling configuration can let a feature map's spatial extent collapse toward 1x1 (or, with an off-by-one padding bug, to a degenerate size) partway through a deep stack; a batch-norm layer sitting downstream of that then computes its variance over a near-constant or degenerate spatial dimension, driving the same near-zero-variance division failure described in step 3 above. When a NaN appears specifically in a CNN, log activation TENSOR SHAPES layer-by-layer, not just tensor values, since a silently-collapsed spatial dimension is easy to miss if you're only watching for NaN values and not for shape degeneracy.
Worked example
A concrete instance of step 3: computing log(softmax(z)) as two separate operations can, for a very confident wrong prediction, produce a softmax output that underflows to EXACTLY 0.0 in floating point (not just a small number), so the subsequent log(0.0) returns -inf, and a -inf value contaminates every downstream gradient computation that touches it. The fused log_softmax operation instead computes this directly via the log-sum-exp identity, logsoftmax(z)i=zi−max(z)−log∑jezj−max(z), which never actually materializes the underflowing intermediate probability, so it cannot produce this failure mode.
Trade-offs & pitfalls
A common mistake is jumping straight to lowering the learning rate and moving on once the NaN stops, without confirming the NaN doesn't recur later at a different step; a learning rate drop can mask a genuine data or numerical-op bug by making it statistically less likely to trigger, rather than eliminating the actual cause. Instrumenting the pipeline to log the loss value, gradient norm, and the first tensor to go non-finite (rather than waiting for the aggregate loss to visibly become NaN several steps later) makes this entire investigation far faster the next time it happens, and framework anomaly-detection hooks (step 8) are cheap enough to leave enabled during any run that has previously shown instability.
Unlock Full Question Bank
Get access to all Deep Learning: Neural Networks and Architectures interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.