Python Programming Questions
Python as an interview language: core syntax, data types and built-in collections, comprehensions, iterators and generators, idiomatic style, and the standard library, extending into data-oriented and automation use of the language and its common libraries. Covers writing correct, Pythonic code and reasoning about the language's semantics. The most heavily exercised language surface in this category across engineering and data roles.
Explain the difference between a list comprehension, generator expression, and using map/filter in Python. When would you prefer each? Give a short code example (2-3 lines) for each showing how to square numbers 0..9.
Sample Answer
Difference & when to prefer
- List comprehension: produces a list eagerly; concise and fast for moderate-sized results when you need random access or repeated iteration.
- Generator expression: lazy iteration, low memory; prefer for large streams or pipeline processing.
- map/filter: functional style; can be slightly faster in some cases and composes well with other functions or builtins; returns iterator in Py3.
Examples (square 0..9):
List comprehension:
squares = [x*x for x in range(10)]
print(squares)
# [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
Generator expression:
squares_gen = (x*x for x in range(10))
print(list(squares_gen))
# [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
map/filter:
squares_map = list(map(lambda x: x*x, range(10)))
print(squares_map)
# [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
All three print the identical list, confirming the three forms are interchangeable here; only their evaluation style (eager list, lazy generator, eager map/filter) differs.
You're designing the public exception types for a library other teams will depend on. When do you define custom exception classes versus reusing built-ins, how narrow should an except clause be, and how do you use exception chaining (raise ... from ...) to preserve the original cause?
Sample Answer
Direct answer
Define a custom exception when a caller needs to programmatically distinguish and handle a specific failure mode; reuse a built-in (ValueError, TypeError, KeyError) when the failure is a generic, well-understood violation with no library-specific handling to offer. Keep except clauses as narrow as the exception you can actually recover from, never bare except:. Use raise NewError(...) from original whenever you translate a low-level exception into a library-level one, so the original traceback and type are preserved for debugging instead of discarded.
Structured elaboration
When to define a custom exception type
- Define one when a caller might reasonably want to catch this specific failure and do something different for it than for other failures (retry, fall back, surface a specific user-facing message). If every caller would handle it the same way as a generic
ValueError, a custom type adds ceremony without adding value. - Root the hierarchy at a single package-level base (for example
MyLibError(Exception)), with specific failures subclassing it (ModelLoadError(MyLibError),InvalidDatasetError(MyLibError)). This lets a caller catch the single base class to mean "anything this library can go wrong in" without having to enumerate every subclass, while still allowing narrower catches where useful. - Name exceptions for what failed semantically, not for the internal mechanism that detected it; a caller should be able to catch
ModelLoadErrorwithout knowing or caring whether the implementation currently reads checkpoints from disk, S3, or a database.
How narrow an except clause should be
- Catch the most specific exception type you can actually do something about. A library function should generally let exceptions it cannot meaningfully handle propagate (or wrap them in its own type), rather than catching broadly and hiding the failure.
- Bare
except:(orexcept Exception:used as a catch-all) at the library level is almost always wrong: it catches things likeKeyboardInterrupt-adjacent control-flow signals in the case of bareexcept:, and in the case ofexcept Exception:it hides programming errors (a typo causing anAttributeError) behind the same handling path as an expected, recoverable failure. - Application-level code (the outermost layer, close to a user or an operator) is where broader catches are more defensible, specifically to provide a fallback, a user-facing error message, or a metric increment, since at that point there is nowhere further up to propagate to.
Exception chaining with raise ... from ...
raise ModelLoadError(...) from original_excsetsoriginal_excas the new exception's__cause__, so the traceback shown to a developer includes both: "the following exception occurred while handling this one," preserving the original type, message, and traceback instead of losing them.- Omitting
from(raise ModelLoadError(...)inside anexceptblock) still implicitly chains the original as__context__, shown as "during handling of the above exception, another exception occurred"; explicitfromis preferred when the translation is intentional, since it documents the relationship as deliberate rather than incidental.raise ... from Nonesuppresses the chain entirely, which is appropriate only when the original exception is genuinely irrelevant noise (rare in a library boundary).
Worked example
class LibraryError(Exception):
'''Base class for all errors raised by this library.'''
class ModelLoadError(LibraryError):
'''Raised when a model checkpoint cannot be loaded.'''
def load_checkpoint(path):
raise FileNotFoundError(path)
try:
load_checkpoint("/tmp/does-not-exist.pt")
except FileNotFoundError as e:
raise ModelLoadError(f"could not load checkpoint at {e}") from e
This raises ModelLoadError: could not load checkpoint at /tmp/does-not-exist.pt, and the traceback CPython prints includes both exceptions, joined by the line "The above exception was the direct cause of the following exception:", with the original FileNotFoundError and its own traceback shown first. A caller who only knows about this library's API can catch ModelLoadError (or the LibraryError base) without needing to know the failure originated from a missing file rather than, say, a corrupted checkpoint format; a developer debugging the failure still sees the full original traceback via the chained __cause__.
Trade-offs & pitfalls
- A hierarchy that is too deep (many single-use subclasses that no caller ever catches individually) adds API surface for no behavioral benefit; a hierarchy that is too flat (one exception type for every failure) forces every caller to parse the message string to distinguish cases, which is fragile and not something the standard library's own conventions encourage.
- Catching broadly "to be safe" inside a library function is the single most common way debugging information gets lost: an unrelated bug (a typo, an off-by-one) gets silently reclassified as the same expected failure the
exceptclause was written for, and the real bug ships unnoticed. - Changing which exception type a public function raises (or removing a subclass from the hierarchy) is a breaking API change for any caller who catches it specifically; treat the exception hierarchy itself as part of the library's versioned public contract, not as an implementation detail.
Design a caching decorator whose entries expire after a configurable TTL. What data structure backs the cache, how do you evict stale entries (lazily on access versus proactively), and what changes if you also need it to be thread-safe?
Sample Answer
Approach
Back the cache with a plain dict: keys are a canonical (args, sorted(kwargs.items())) tuple, values are (result, expires_at) pairs, and expires_at is measured with time.monotonic() rather than time.time() so the TTL is not affected by the system clock being adjusted (NTP sync, manual clock changes) mid-run. Eviction happens lazily: every access checks whether its own entry is stale and drops it if so, which means a key that is never looked up again just sits in memory until the janitor (below) or a fresh call for that key clears it. Adding thread-safety is a single threading.Lock guarding every read and write of the dict, since dict operations are not safe to interleave with a delete.
Code (Python 3.12)
import time
import threading
from functools import wraps
def ttl_cache(ttl_seconds):
'''Decorator factory: cache each call's result for ttl_seconds.
Thread-safe: one lock guards every read and write of the cache.
'''
def decorator(func):
store = {} # key -> (value, expires_at)
lock = threading.Lock()
@wraps(func)
def wrapper(*args, **kwargs):
key = (args, tuple(sorted(kwargs.items())))
now = time.monotonic()
with lock:
cached = store.get(key)
if cached is not None:
value, expires_at = cached
if expires_at > now:
return value
del store[key] # lazy eviction: stale, drop it
value = func(*args, **kwargs)
store[key] = (value, now + ttl_seconds)
return value
def cache_clear():
with lock:
store.clear()
wrapper.cache_clear = cache_clear
return wrapper
return decorator
calls = {"n": 0}
@ttl_cache(ttl_seconds=0.05)
def slow_square(x):
calls["n"] += 1
return x * x
print(slow_square(4)) # 16
print(slow_square(4)) # 16, served from cache
print(calls["n"]) # 1
time.sleep(0.06)
print(slow_square(4)) # 16, recomputed: previous entry had expired
print(calls["n"]) # 2
Key points
- A dict gives average O(1) lookup by (args, kwargs), which is the right structure whenever the key space is a finite, hashable set of call signatures; if an argument is unhashable (a
list, adict), building the key itself raisesTypeErrorbefore the cache is even consulted. - Lazy eviction (check-and-drop on access) costs nothing extra for keys that keep getting called, but a key that is called once and never again just occupies memory forever, since nothing ever revisits it to notice it went stale. A proactive sweep fixes that: a
daemon=Truebackground thread that wakes up on an interval, takes the lock, and removes every entry whoseexpires_athas passed, independent of whether anyone accesses it. The two are complementary, not either/or: lazy handles the common case cheaply, proactive bounds worst-case memory for cold keys. - Thread-safety changes exactly one thing structurally: every dict read-then-maybe-write sequence in
wrapperhas to happen atomically with respect to other threads, or two threads can both see a miss, both callfunc, and one result silently overwrites the other (wasted work, not corruption, since both results are equally valid, but still wrong for a function with side effects). A single lock around the whole "check, maybe compute, store" block is sufficient here because the cached function's own body is not being timed for lock duration; iffuncitself is slow, holding the lock across the call tofuncserializes unrelated cache misses on different keys too, which is the real cost of this simple approach: it trades throughput under contention for a small, easy-to-reason-about critical section. A production version would narrow the lock to just the dict operations and use a per-key lock (or anif key not in store: store[key] = SENTINELdouble-checked pattern) to let different keys compute concurrently.
Complexity
Per call: O(1) average for the dict lookup/insert (hashing the key), plus whatever func itself costs on a miss. Space: O(u) where u is the number of distinct, not-yet-expired argument combinations currently cached.
Edge cases
- Unhashable arguments raise
TypeErrorwhen the key tuple is built, beforefuncever runs; this is the same failure mode as any dict-keyed cache and is not TTL-specific. - A function with side effects (writes to a file, mutates a global) should generally not be cached at all, since a cache hit silently skips the side effect on the second call.
- Two related, absorbed variants of this same design:
- "Time the call and cache it": extend the stored tuple to
(result, expires_at, duration), timingfuncwithtime.perf_counter()around the call on a miss, so callers can inspect how expensive an entry was to produce (useful for deciding TTL length empirically) without changing the cache's core structure. - "Cache to disk via pickle": persisting
storeacross process restarts withpickle.dump/pickle.loadtrades in-memory-only simplicity for durability, at the cost of needing every cached value (and every key, since tuples of primitives pickle fine but arbitrary objects may not) to be picklable, plus a decision about whether a pickled entry'sexpires_atshould still be honored after the process was down for a while (it should: monotonic time does not survive a restart, so persisted entries need to be re-validated against wall-clock-derived expiry, or simply treated as expired on load).
- "Time the call and cache it": extend the stored tuple to
You need to package a Python library used in data science which includes compiled C extensions and optional GPU support. Outline a cross-platform build and distribution strategy that simplifies installation for users on Linux, macOS, and Windows. Include CI steps and how you'd support pip installs.
Sample Answer
What a wheel is, first: a wheel (file extension .whl) is a pre-built, ready-to-install package file, so pip install mypkg just unpacks it, with no compiler or build step required on the user's own machine. Without wheels, pip has to compile the package's C extensions locally every time, which is slow and fails constantly on machines that lack the right compiler and headers. A concrete trace of the whole flow: a user on 64-bit Linux running Python 3.12 types pip install mypkg[cuda]; pip resolves the closest matching wheel filename it can find on PyPI, something like mypkg-1.0.0-cp312-cp312-manylinux2014_x86_64.whl, downloads it, and installs it directly with no compilation at all. manylinux2014 in that filename is a compatibility tag: it tells pip this wheel was built against an old-enough baseline of Linux system libraries that it will run correctly on nearly any modern Linux distribution, not just the exact one it was built on.
Strategy: provide manylinux and macOS/Windows wheels for common Python versions; fall back to source+build for edge cases. Offer GPU optional extras (extra tags like package[cuda]).
Build & CI:
- Linux: use GitHub Actions with cibuildwheel (a CI tool that automates building a correctly-tagged, pip-installable wheel for every OS/Python-version combination you need, instead of hand-writing that build matrix yourself) to produce manylinux2014 wheels for x86_64 and aarch64; build both CPU and GPU wheels (GPU via CUDA toolkits in separate matrix jobs).
- macOS: use cibuildwheel to build macOS universal2 wheels (a single wheel file containing binaries for both Intel and Apple Silicon Macs, so pip does not need two separate mac wheel variants).
- Windows: use cibuildwheel on windows-latest to build wheels.
- Run tests in each job, run integration tests with/without GPU.
Packaging:
- Use cython/setuptools or scikit-build with CMake (scikit-build bridges Python's packaging tools with CMake, the standard build system for compiled C/C++ code, so the compiled extension gets built the same way on every platform) for portability. Produce wheels with bundled libs where license allows.
- Publish wheels to PyPI and use tags: mypkg, mypkg[cuda] ==> extra_requires (Python packaging's mechanism for optional, named dependency sets: a user who runs
pip install mypkg[cuda]pulls in the extra CUDA-related dependencies, while a plainpip install mypkgskips them) that pin appropriate CUDA runtime and optional GPU wheels.
User installs:
- pip will pick wheel; if none, pip falls back to build from source (document build deps). Provide conda-forge (a large, community-maintained collection of prebuilt Conda packages) packages for easier GPU/runtime management.
Notes: sign wheels, provide concise install docs, and CI nightly builds for new Python/OS combos. If you can only do one thing first, do this: get cibuildwheel producing manylinux2014 plus macOS plus Windows CPU wheels through GitHub Actions and publish them to PyPI, since that alone lets the overwhelming majority of users pip install mypkg with zero compiler needed; GPU-specific wheels, conda-forge packages, and wheel signing are all valuable follow-ups, not blockers to a usable first release.
Write a function that converts a messy numeric string, things like '1,234.56', '$1.2M', 'NaN', or an empty string, into a float, returning None for anything unparseable. Would you check the input's shape before converting (look before you leap) or just try the conversion and catch the exception (ask forgiveness)? Justify your choice here.
Sample Answer
Approach
Ask forgiveness, not permission: normalize away the known messy formatting (thousands separators, a currency symbol, a K/M/B magnitude suffix, parenthesized negatives), then attempt the actual conversion with float() inside a try/except ValueError, returning None on failure. Look-before-you-leap (LBYL) would mean writing a regex or a hand-rolled validator that decides in advance whether the string looks convertible, which in practice means re-implementing everything float()'s own parser already does correctly, just to decide whether to call it; that duplicated logic is extra surface area that can drift from what float() actually accepts, and it does not save the try/except anyway, since malformed input mixing multiple messy features ("$1,2M3") will still need a runtime check. EAFP here means: only do the normalization steps that resolve known, named messy formats, then let Python's own float parser be the single source of truth on whether the result is actually a valid number.
Code (Python 3.12)
import re
_SUFFIX_MULTIPLIER = {"K": 1e3, "M": 1e6, "B": 1e9}
_CURRENCY_CHARS = re.compile(r"[$,\s]")
def parse_number(raw):
'''Convert a messy numeric string to float; return None if unparseable.
Handles thousands separators (1,234.56), a leading currency symbol ($),
a K/M/B magnitude suffix, and parenthesized negatives ((3.5K) == -3500).
Treats 'nan' and '' as missing data (None), not as float('nan'): in this
pipeline a NaN token means "value absent", not "the value is not-a-number".
'''
if raw is None or not isinstance(raw, str):
return None
text = raw.strip()
if not text or text.lower() == "nan":
return None
negative = text.startswith("(") and text.endswith(")")
if negative:
text = text[1:-1]
multiplier = 1.0
if text and text[-1].upper() in _SUFFIX_MULTIPLIER:
multiplier = _SUFFIX_MULTIPLIER[text[-1].upper()]
text = text[:-1]
text = _CURRENCY_CHARS.sub("", text)
try:
value = float(text)
except ValueError:
return None
value *= multiplier
return -value if negative else value
cases = ["1,234.56", "$1.2M", "NaN", "", " 42 ", "(3.5K)", "abc", None, "3.1e2"]
for c in cases:
print(repr(c), "->", parse_number(c))
# Variant: average a batch of messy values while ignoring missing ones.
def average_ignoring_missing(raw_values):
parsed = [v for v in (parse_number(x) for x in raw_values) if v is not None]
return sum(parsed) / len(parsed) if parsed else None
print(average_ignoring_missing(["10", "", "20", "NaN", "30"]))
Output:
'1,234.56' -> 1234.56
'$1.2M' -> 1200000.0
'NaN' -> None
'' -> None
' 42 ' -> 42.0
'(3.5K)' -> -3500.0
'abc' -> None
None -> None
'3.1e2' -> 310.0
20.0
Key points
- Every normalization step (strip whitespace, strip a currency symbol, strip a magnitude suffix, strip parentheses) is applied unconditionally and cheaply; only the final
float(text)call is wrapped intry/except, so the EAFP boundary is drawn as tightly as possible around the one operation whose success genuinely cannot be known in advance without duplicating its logic. "NaN"is deliberately treated asNone(missing), not passed through tofloat("nan")(which would succeed and produce an actual NaN float): the two are different data-quality signals, "this field was never populated" versus "this field holds the floating-point value NaN," and conflating them would let a NaN silently poison a downstreamsum()or comparison (nan != nanin every comparison, including equality, which produces confusing bugs if it leaks into arithmetic unexpectedly).isinstance(raw, str)up front is LBYL for exactly one thing: rejecting non-string input (alist, anintalready) before doing string operations on it that would raiseAttributeError/TypeErrorrather than theValueErrorthis function is designed to swallow. Mixing that one type check into an otherwise EAFP function is a normal, common pattern, not a contradiction: LBYL and EAFP are not mutually exclusive within a single function, the choice is made per failure mode.
Complexity and edge cases
O(k) where k is the length of the string, dominated by the regex substitution and float()'s own parsing.
- Whitespace-only strings (
" ") become""after.strip()and correctly returnNone. - A currency symbol combined with a suffix and a thousands separator all at once (
"$1,234.5K") is handled correctly because the three normalization steps are independent and composable. - Variant: averaging while ignoring missing values (shown in the code above via
average_ignoring_missing): parse every value with the sameparse_number, filter out theNones, then average what remains, so a missing reading never silently counts as zero and never crashes the average. - Variant: a safe money sum:
floataccumulates binary floating-point rounding error across many additions, which is unacceptable for currency; the EAFP shape stays the same but the target type changes todecimal.Decimal, which represents decimal fractions exactly:
import re
from decimal import Decimal, InvalidOperation
_CURRENCY_CHARS = re.compile(r"[$,\s]")
def parse_money(raw):
'''Like parse_number but returns Decimal for exact currency arithmetic.'''
if raw is None or not isinstance(raw, str):
return None
text = _CURRENCY_CHARS.sub("", raw.strip())
if not text:
return None
try:
return Decimal(text)
except InvalidOperation:
return None
prices = ["$19.99", "$5.01"]
total = sum((parse_money(p) for p in prices), start=Decimal("0"))
print(total) # 25.00, exact
Unlock Full Question Bank
Get access to all Python Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.