Python Programming Questions
Python as an interview language: core syntax, data types and built-in collections, comprehensions, iterators and generators, idiomatic style, and the standard library, extending into data-oriented and automation use of the language and its common libraries. Covers writing correct, Pythonic code and reasoning about the language's semantics. The most heavily exercised language surface in this category across engineering and data roles.
Given a Pandas DataFrame df with columns ['user_id', 'event_time', 'value'], write idiomatic code to compute for each user the rolling 7-day sum of 'value' based on event_time (which is a datetime). Ensure the solution scales for millions of rows.
Sample Answer
Requirements: per-user 7-day rolling sum by event_time, scalable.
Idiomatic, vectorized Pandas (vectorized meaning the loop over rows runs as fast, compiled bulk code instead of a slow Python-level for loop, one row at a time):
import pandas as pd
df = pd.DataFrame({
"user_id": ["u1", "u1", "u1", "u1", "u2"],
"event_time": pd.to_datetime([
"2024-01-01", "2024-01-03", "2024-01-05", "2024-01-09", "2024-01-02"
]),
"value": [10, 20, 30, 40, 5],
})
# ensure datetime and sort
df['event_time'] = pd.to_datetime(df['event_time'])
df = df.sort_values(['user_id', 'event_time'])
# set index for time-based rolling and compute 7-day sum per user
result = (
df.set_index('event_time')
.groupby('user_id')['value']
.rolling('7D')
.sum()
.reset_index(name='rolling_7d_sum')
)
Worked example, verified with pandas on CPython 3.12:
df = pd.DataFrame({
"user_id": ["u1", "u1", "u1", "u1", "u2"],
"event_time": pd.to_datetime([
"2024-01-01", "2024-01-03", "2024-01-05", "2024-01-09", "2024-01-02"
]),
"value": [10, 20, 30, 40, 5],
})
print(result) # (built by running df through the code above)
Output:
user_id event_time rolling_7d_sum
0 u1 2024-01-01 10.0
1 u1 2024-01-03 30.0
2 u1 2024-01-05 60.0
3 u1 2024-01-09 90.0
4 u2 2024-01-02 5.0
Tracing u1's rows: Jan 1 has no prior rows in its trailing 7-day window, so the sum is just its own value, 10. Jan 3's window (Jan 3 back to Dec 27) includes Jan 1 and Jan 3, 10 + 20 = 30. Jan 5's window (back to Dec 29) includes Jan 1, 3, and 5, 10 + 20 + 30 = 60. Jan 9's window (back to Jan 2) EXCLUDES Jan 1, which is 8 days earlier, outside the 7-day cutoff, so it sums only Jan 3, 5, and 9: 20 + 30 + 40 = 90. This is exactly what makes it a time-based, not row-count-based, window: the number of prior rows included varies depending on how many actually fall within the trailing 7 real days, not a fixed count of the last N rows.
Notes: Uses time-based rolling with groupby which is vectorized and memory-efficient for large data. For millions of rows, ensure event_time is datetime64, operate on chunked parquet files (parquet: a compressed, columnar file format for tabular data, well suited to reading only the columns and row-groups a job actually needs) if memory constrained, and consider Dask or PySpark for distributed scaling once the data genuinely no longer fits on one machine: Dask mirrors the Pandas API but splits data into partitions and runs the same groupby/rolling-style operations across them in parallel; PySpark is a different, JVM-backed distributed engine with its own (similar but not identical) DataFrame API. Plain, chunked Pandas is enough for anything that still fits on a single machine's memory; reach for Dask/PySpark only once it genuinely does not.
Describe how Python's reference counting and garbage collector work together to reclaim memory. Provide an example of an object pattern that requires the garbage collector (i.e., not reclaimed by reference counting alone).
Sample Answer
How they work together:
- CPython uses reference counting: each object tracks references; when count hits zero it's reclaimed immediately.
- The cyclic GC (generational, meaning it groups objects by how long they've survived so far and scans the youngest group most often, since most garbage dies young and rarely needs a second look) complements this by detecting reference cycles that reference counting alone can't free.
Example needing GC:
Two objects referencing each other (cycle) with del avoided or handled by GC. Example:
class Node:
def __init__(self):
self.ref = None
a = Node()
b = Node()
a.ref = b
b.ref = a
# If no external refs remain, ref counts are non-zero due to cycle -> GC must collect
Actually watching the reclamation happen, verified on CPython 3.12, extending the example above with a __del__ and an explicit gc.collect():
import gc
class Node:
def __init__(self, name):
self.name = name
self.ref = None
def __del__(self):
print(f'__del__ called for {self.name}')
a = Node('a')
b = Node('b')
a.ref = b
b.ref = a
del a, b # neither refcount reaches zero; both still reference each other
gc.collect() # the cyclic GC finds the unreachable pair and reclaims it
Output:
__del__ called for a
__del__ called for b
Both finalizers ran as part of gc.collect() reclaiming the cycle, confirming the cycle really was found and cleaned up, not just theoretically collectible.
Reference counting handles most deallocations quickly; the cyclic GC periodically finds unreachable cycles and reclaims them. Cycles involving objects with __del__ used to require special handling: on Python versions before 3.4, the cyclic GC could not safely decide what order to finalize such objects in, so it left them permanently uncollected in gc.garbage instead of guessing. Since Python 3.4 (PEP 442), __del__ is called safely as part of collecting the cycle, exactly as demonstrated above, and the objects are then reclaimed normally; citing the old "cycles with __del__ can't be collected" claim as current fact is a dated answer today.
A process automation tool needs to validate hundreds of files in parallel, but the final summary must be emitted in the same order the files were submitted. A fatal parse error should stop remaining work as quickly as possible. How would you structure the goroutines, communication, and shutdown logic in Go?
Sample Answer
Go vocabulary, translated for a Python reader: a goroutine is Go's lightweight, concurrently-running function, similar in spirit to a Python thread but far cheaper to start and typically used in much greater numbers in real Go code; a channel is a typed, thread-safe queue you send values into and receive values out of, roughly like a queue.Queue shared between Python threads, except the compiler enforces the type of what flows through it; a WaitGroup is a counter that lets the caller block until every worker goroutine has signaled it is done, similar to calling .join() on a list of Python Thread objects; ctx (short for context) carries a shared cancellation signal through the call tree, and ctx.Done() returns a channel that closes the moment that signal fires, so any goroutine can cheaply check "has someone asked everything to stop?" without polling a shared boolean, comparable to checking a Python threading.Event.
Structure
I’d use a bounded worker pool. A feeder sends files with an index into a jobs channel. Each worker validates one file, sends {index, result, err} to a results channel, and watches ctx.Done() so context cancellation, meaning a cooperative stop signal, is fast.
Ordering
A single collector keeps nextIndex and a map of out-of-order results. If result 7 arrives before 6, store it until 6 is ready, then flush in submission order.
Fatal parse error
If a worker sees an unrecoverable parse error, it sends the error and calls cancel(). That stops the feeder, makes workers exit on ctx.Done(), and prevents new work from starting.
Shutdown
- close
jobsafter feeding stops WaitGroupwaits for workers- close
resultsafter workers finish - collector drains until closed or canceled
Worked example: files 1, 2, 3, 4 arrive. If file 3 has a fatal parse error, 1 and 2 can still be emitted, 4 is never started, and the summary reports the error immediately.
Shape of the code (a sketch of the structure, not a full compiled program):
type job struct {
index int
path string
}
type result struct {
index int
output string
err error
}
func run(ctx context.Context, paths []string, numWorkers int) []result {
ctx, cancel := context.WithCancel(ctx)
defer cancel()
jobs := make(chan job)
results := make(chan result)
var wg sync.WaitGroup
for w := 0; w < numWorkers; w++ {
wg.Add(1)
go func() {
defer wg.Done()
for j := range jobs {
out, err := validate(j.path)
if err != nil && isFatal(err) {
cancel()
}
results <- result{index: j.index, output: out, err: err}
}
}()
}
go func() {
for i, p := range paths {
select {
case jobs <- job{index: i, path: p}:
case <-ctx.Done():
close(jobs)
return
}
}
close(jobs)
}()
go func() {
wg.Wait()
close(results)
}()
return collectInOrder(results)
}
The jobs and results lines are the channel declarations; the three go func() { ... }() blocks are the goroutine launches, one pool of workers, one feeder, one closer. select { case jobs <- job{...}: ... case <-ctx.Done(): ... } is how the feeder stays responsive to cancellation even while trying to send a job that a worker isn't ready to receive yet, instead of blocking forever on a full channel after a fatal error.
This gives parallelism, ordered output, and fast failure without deadlocks.
You need to package a Python library used in data science which includes compiled C extensions and optional GPU support. Outline a cross-platform build and distribution strategy that simplifies installation for users on Linux, macOS, and Windows. Include CI steps and how you'd support pip installs.
Sample Answer
What a wheel is, first: a wheel (file extension .whl) is a pre-built, ready-to-install package file, so pip install mypkg just unpacks it, with no compiler or build step required on the user's own machine. Without wheels, pip has to compile the package's C extensions locally every time, which is slow and fails constantly on machines that lack the right compiler and headers. A concrete trace of the whole flow: a user on 64-bit Linux running Python 3.12 types pip install mypkg[cuda]; pip resolves the closest matching wheel filename it can find on PyPI, something like mypkg-1.0.0-cp312-cp312-manylinux2014_x86_64.whl, downloads it, and installs it directly with no compilation at all. manylinux2014 in that filename is a compatibility tag: it tells pip this wheel was built against an old-enough baseline of Linux system libraries that it will run correctly on nearly any modern Linux distribution, not just the exact one it was built on.
Strategy: provide manylinux and macOS/Windows wheels for common Python versions; fall back to source+build for edge cases. Offer GPU optional extras (extra tags like package[cuda]).
Build & CI:
- Linux: use GitHub Actions with cibuildwheel (a CI tool that automates building a correctly-tagged, pip-installable wheel for every OS/Python-version combination you need, instead of hand-writing that build matrix yourself) to produce manylinux2014 wheels for x86_64 and aarch64; build both CPU and GPU wheels (GPU via CUDA toolkits in separate matrix jobs).
- macOS: use cibuildwheel to build macOS universal2 wheels (a single wheel file containing binaries for both Intel and Apple Silicon Macs, so pip does not need two separate mac wheel variants).
- Windows: use cibuildwheel on windows-latest to build wheels.
- Run tests in each job, run integration tests with/without GPU.
Packaging:
- Use cython/setuptools or scikit-build with CMake (scikit-build bridges Python's packaging tools with CMake, the standard build system for compiled C/C++ code, so the compiled extension gets built the same way on every platform) for portability. Produce wheels with bundled libs where license allows.
- Publish wheels to PyPI and use tags: mypkg, mypkg[cuda] ==> extra_requires (Python packaging's mechanism for optional, named dependency sets: a user who runs
pip install mypkg[cuda]pulls in the extra CUDA-related dependencies, while a plainpip install mypkgskips them) that pin appropriate CUDA runtime and optional GPU wheels.
User installs:
- pip will pick wheel; if none, pip falls back to build from source (document build deps). Provide conda-forge (a large, community-maintained collection of prebuilt Conda packages) packages for easier GPU/runtime management.
Notes: sign wheels, provide concise install docs, and CI nightly builds for new Python/OS combos. If you can only do one thing first, do this: get cibuildwheel producing manylinux2014 plus macOS plus Windows CPU wheels through GitHub Actions and publish them to PyPI, since that alone lets the overwhelming majority of users pip install mypkg with zero compiler needed; GPU-specific wheels, conda-forge packages, and wheel signing are all valuable follow-ups, not blockers to a usable first release.
Design a small library API in Python for vectorized string transformations on large Pandas Series that avoids creating multiple temporaries for chained operations (e.g., s.str.lower().str.replace(...).str.strip()). Sketch API and explain implementation strategies to minimize allocations.
Sample Answer
Requirements & idea: Provide a lazy, composable API that records string ops and applies them in a single pass to avoid temporaries. Offer a lightweight proxy object wrapping Series with an operation pipeline executed in-place or chunked.
API sketch, with a working _apply_pipeline (the earlier sketch left this function unimplemented; here it actually runs each recorded op vectorized, once, over the whole chunk, rather than materializing an intermediate Series between every step):
class StrChain:
def __init__(self, series):
self.series = series
self.ops = []
def lower(self):
self.ops.append(('lower', None)); return self
def replace(self, pat, repl):
self.ops.append(('replace', (pat, repl))); return self
def strip(self):
self.ops.append(('strip', None)); return self
def compute(self, chunk_size=10_000):
# apply ops chunk-wise to avoid temporaries
return _apply_pipeline(self.series, self.ops, chunk_size)
def _apply_pipeline(series, ops, chunk_size):
parts = []
for start in range(0, len(series), chunk_size):
chunk = series.iloc[start:start + chunk_size]
for op_name, arg in ops:
if op_name == 'lower':
chunk = chunk.str.lower()
elif op_name == 'replace':
pat, repl = arg
chunk = chunk.str.replace(pat, repl, regex=False)
elif op_name == 'strip':
chunk = chunk.str.strip()
parts.append(chunk)
return type(series)(pd.concat(parts)) if parts else series
_apply_pipeline walks the recorded op list once per chunk (default: the whole Series in one chunk, for anything that fits in memory), calling the real vectorized Series.str method for each recorded op in sequence; "chunk-wise" here means each chunk only ever holds one intermediate Series at a time (reassigned to chunk on each step) rather than every stage's output existing simultaneously the way s.str.lower().str.replace(...).str.strip() chained directly would briefly do.
Worked example, verified with pandas on CPython 3.12:
import pandas as pd
s = pd.Series([' Hello World ', ' FOO-BAR ', ' Already lower '])
result = StrChain(s).lower().replace('-', ' ').strip().compute()
print(list(result))
Output:
['hello world', 'foo bar', 'already lower']
Each string is lowercased, has - replaced with a space, and is stripped of surrounding whitespace, in that recorded order, confirming the chain actually runs end to end and produces the same result s.str.lower().str.replace('-', ' ', regex=False).str.strip() would, just without materializing three separate full-Series temporaries to get there.
Implementation strategies:
- Represent ops as vectorized functions (use Series.str methods or numpy.char).
- Apply pipeline per chunk: read chunk, apply all ops in sequence in-place (reuse buffers), write out to result array or new Series with preallocated dtype.
- For unicode/regex heavy ops, compile regex ahead.
- Use numba (compiles a plain Python function to machine code the first time it runs) or cython (compiles Python-like code all the way down to C, with explicit type declarations, for the most control and the largest potential speedup) for the hot path once profiling shows plain vectorized
Series.strops are the actual bottleneck; for most chains, the vectorized ops above are already enough and neither is needed by default.
Minimize allocations:
- Reuse a single buffer per chunk; preallocate numpy object/bytes arrays when possible.
- Fuse operations into one pass (e.g., lower+strip -> single routine, implemented as one
numpy.charor Cython function operating character-by-character instead of two separate full passes over the data).
Trade-offs: chunking adds overhead but limits peak memory; fusing ops increases implementation complexity.
Unlock Full Question Bank
Get access to all 46 Python Programming interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.