Data Ingestion and Source System Integration Questions
Getting data out of heterogeneous source systems and landing it reliably: APIs, operational databases, file drops, webhooks, message queues and third-party SaaS. Covers connector selection and design (managed platforms versus Debezium, DMS or Kafka Connect versus building your own), pull versus push and polling versus webhook patterns, incremental extraction and high-watermark strategy including what to do when a source offers no native change capture, authentication and credential rotation against third-party APIs, source-side rate limits and quotas, schema drift and contract breakage at the source boundary, backfill and replay of history, ingestion-time data-quality gates, reconciliation after a source outage, and negotiating with source-system owners. The scope stops at the boundary: once data has landed, transforming it, the architecture of the pipeline that carries it, stream-processing mechanics, and pipeline monitoring are all covered separately.
Tell me about a time you worked directly with the owners of a source system to reduce the operational impact your ingestion pipeline had on them, for example cutting the lock contention or CPU load your extraction was putting on their transactional database. What was the situation, what did you propose, what actually changed, and what did you learn about working with a team that owns data you depend on but do not control?
Sample Answer
Direct answer
I once inherited an hourly extraction job pulling from a transactional order database that, on close inspection, was holding row locks for several seconds per batch during business hours, occasionally causing brief but visible latency spikes for the application team's own customer-facing checkout flow. I worked with the application team to move the extraction to a read replica and switch from a naive full-table scan to a keyset-based incremental query, which eliminated the lock contention entirely, and the collaboration itself taught me that a source team's trust, once damaged, is much harder to rebuild than the technical fix itself.
Structured elaboration
The situation
- The application team flagged intermittent checkout latency spikes that correlated, timing-wise, with our hourly extraction job, though nobody on either team had proven the connection yet.
- Investigating the extraction query showed it was doing a full-table scan with an implicit lock-heavy read pattern, on the primary database, during business hours, not the off-peak window it had originally been designed for years earlier.
What I proposed
- Move extraction to the existing read replica, so the query's I/O and lock impact would no longer touch the primary at all.
- Rewrite the query from a full-table scan to a keyset-based incremental pull (using the same reliable timestamp column already available), which reduced the actual amount of data read per run dramatically on top of removing it from the primary entirely.
- I proposed this jointly, in a working session with the application team's on-call engineer, rather than unilaterally deciding and just informing them afterward, specifically because they were the ones who would bear the risk if my read-replica assumption about staleness turned out to be wrong for their use case.
What actually changed
- The extraction now runs against the replica on the same hourly schedule, with the incremental rewrite cutting the query's own runtime by roughly 80% as a side benefit.
- The application team's own monitoring confirmed the checkout latency spikes disappeared entirely once the change shipped, giving both teams independent confirmation rather than just my own before-and-after measurement.
What I learned
- The technical fix (read replica plus a better query) was genuinely the easy part; the harder part was that the application team had, reasonably, become somewhat guarded about any request touching "their" database after a previous unrelated incident, and earning enough trust to even get time on their calendar to discuss the problem took more deliberate relationship-building than the fix itself.
- I now default to proactively asking a new source team what their own operational constraints and pain points are BEFORE building an extraction against their system, rather than waiting for a problem to surface and then reacting to it, since that first framing sets the tone for the whole ongoing relationship.
Worked example
A structurally similar situation, negotiating an event contract with an application team before ingestion even began, followed the same underlying pattern: rather than reverse-engineering their event schema from whatever happened to be emitted, I asked for sample payloads, the expected field set and types, and realistic volume estimates up front, in a short working session before any pipeline code was written. When the application team later needed to add a field mid-project, having already established that channel meant they proactively looped me in on the schema change before shipping it, rather than my discovering it after the fact from a broken downstream job, which is the outcome the same underlying discipline, treating the source-owning team as an ongoing partner rather than a one-time data tap, was meant to produce.
Trade-offs & pitfalls
- Proposing a technically superior fix without first understanding why the source team might be cautious about changes to their system, given their own history and incidents, risks the proposal being received as tone-deaf even when it is technically correct.
- Moving extraction to a read replica is not free: it introduces a small amount of replication lag that has to be an acceptable trade-off for the specific freshness needs of the pipeline, and that trade-off should be made explicitly with the consuming team, not silently assumed.
- A single successful negotiation does not guarantee the relationship stays healthy; ongoing communication (a shared channel, a habit of flagging planned changes ahead of time) matters more over the long run than any one well-handled incident.
- Do not treat "the source team agreed" as the end of the story; verify the fix's actual effect independently (their own monitoring, in this case) rather than relying solely on your own before-and-after measurement, which can be subject to the same blind spots that missed the original problem.
A new source system supports both webhooks and a polling API. Walk through the trade-offs of using webhooks versus periodic API polling for ingesting from it: reliability, retry handling, back-pressure, security, and the operational monitoring each approach requires.
Sample Answer
Direct answer
Webhooks give you near-real-time delivery at the cost of needing a reliably-available receiver and inheriting whatever retry behavior the source implements; polling gives you full control over pacing and a simpler operational model at the cost of latency bounded by your polling interval. When the source supports webhooks, prefer them for anything with a real freshness requirement, but keep a periodic reconciliation poll running alongside as a safety net, since even a well-implemented webhook sender can silently fail to deliver.
Structured elaboration
Reliability
- A webhook depends entirely on the source's own retry policy for delivery; some sources retry aggressively with backoff over hours, some retry a few times and give up, and a few offer no redelivery at all, so "reliability" here is really a question about the SOURCE, not something you control.
- A poll is self-healing by construction: if one poll fails or is missed, the next scheduled poll picks up everything that changed since the last successful checkpoint, with no dependency on the source retrying anything.
Retry handling
- For webhooks, your receiver has to be idempotent, since the source's retries mean you will occasionally receive the same event twice; deduplicate on the event's own ID before processing.
- For polling, retries are entirely under your control: a failed poll is simply retried by your own scheduler with your own backoff policy.
Back-pressure
- A webhook sender pushes at whatever rate events occur, which can spike unpredictably (a bulk operation on the source side, for example); your receiver has to absorb bursts without falling over, typically by accepting quickly and queuing for async processing rather than processing synchronously in the request handler.
- Polling gives you back-pressure for free: you simply do not ask for more until you are ready.
Security
- A webhook receiver is a public-facing endpoint, so it needs signature verification (an HMAC, a hash-based message authentication code, that the source includes, checked against a shared secret) to confirm a request genuinely came from the source and was not forged.
- Polling has no equivalent public-surface risk, since you are the one initiating every request, but it does still need the outbound credentials properly secured.
Operational monitoring
- For webhooks, monitor for SILENCE, not just errors: a source that has quietly stopped sending events (a misconfiguration on their side, a subscription that expired) produces no errors at all on your end, only an absence of expected traffic, which is much harder to notice than an explicit failure.
- For polling, monitor the poll's own success rate and the volume of records returned per poll, watching for an unexpected drop that would suggest the source-side query or filter has broken.
Worked example
A source offers both a webhook and a polling API for the same event type. Freshness matters (downstream alerts should fire within seconds), so the primary path is the webhook: the receiver verifies the HMAC signature, deduplicates on event ID against a short-TTL store, and queues the event for async processing so a burst of webhooks does not block the HTTP response. Running alongside, a lightweight poll every 15 minutes checks for any records the webhook path might have missed, comparing against the same downstream store; any record present in the source's poll response but absent downstream triggers a targeted backfill of just that gap, which is far cheaper than either abandoning webhooks (losing real-time freshness) or trusting them blindly (risking silent, undetected gaps).
Trade-offs & pitfalls
- If the source offers no webhook at all, the framing is not really "webhook versus polling," it becomes "how tight a poll interval can we sustain without exceeding the source's rate limit," and any real-time expectation needs to be reset against that constraint honestly rather than pretending polling can match webhook latency.
- A webhook receiver that is down even briefly can lose events permanently if the source does not retry, or does not retry for long enough to cover your downtime; know the source's specific retry window before relying on it as your only path.
- Signature verification is frequently skipped in early implementations "to get something working" and then never added later; treat it as non-negotiable for any public-facing receiver, not an enhancement to add eventually.
- The reconciliation-poll safety net is cheap insurance against silent webhook failures, but only if someone is actually watching its output; an unmonitored reconciliation job that has been silently finding and ignoring gaps for months is barely better than having no safety net at all.
You are ingesting data from multiple third-party APIs that use OAuth2 and rotating API keys. Describe how you would securely store and refresh credentials, handle a token-refresh failure without losing data, enforce each source's rate limits, and design retry and backoff so ingestion stays reliable and auditable.
Sample Answer
Direct answer
Credential management for many third-party connectors comes down to three disciplines: store every credential in a secrets manager and never in connector configuration files, refresh OAuth2 tokens proactively before they expire rather than reactively after a call fails, and treat a refresh failure as a distinct, alertable condition rather than letting it silently degrade into skipped syncs. Rate limits and retry behavior then have to be tracked per source, since every third-party API enforces its own limit differently.
Structured elaboration
Secure storage
- Store client IDs, client secrets, and refresh tokens in a dedicated secrets manager (a cloud provider's secrets service or an equivalent), never in a connector's plain configuration file or in version control.
- Scope access narrowly: the connector process should have permission to read only the credentials it needs, not every secret in the organization's store, so a compromised connector cannot pivot to unrelated systems.
Refreshing tokens without losing data
- Refresh proactively, ahead of expiry (for example, when a token has less than 10 minutes of validity left), rather than waiting for a call to fail with a 401 and refreshing reactively; this avoids losing an in-flight batch to an expired token mid-pull.
- If a refresh does fail (the refresh token itself is invalid or revoked), do not let ingestion silently stop: raise a distinct, named alert ("source X requires re-authorization") rather than letting the connector fail the same generic way it would for a transient network error, since these need different human responses.
- Persist enough state that a connector recovering from a refresh failure resumes from its last successful checkpoint rather than needing a full re-pull once access is restored.
Enforcing each source's rate limit
- Each third-party API documents its own limit differently (requests per second, per minute, per day, sometimes per specific endpoint); track each source's limit as explicit configuration rather than a single global assumption.
- A token-bucket (a counter that starts at some capacity, refills by a fixed amount on a fixed schedule such as once per second, and is spent one token per request, so you physically cannot send faster than the refill rate once the initial balance is used up) or sliding-window (which instead counts how many requests actually landed in the trailing N-second window and blocks new ones once that count hits the cap) limiter per source, refilled or evaluated at that source's documented rate, keeps you comfortably under the cap without needing to guess a safe interval empirically.
Retry and backoff, and making it auditable
- Exponential backoff with jitter on 429s and 5xxs, capped so a persistent failure surfaces as an alert instead of retrying forever.
- Log every token refresh, every rate-limit-triggered backoff, and every retry with enough context (which source, which credential, timestamp) that an audit can reconstruct exactly what happened to a given source's access over time, which matters both for debugging and for satisfying a security review.
Worked example
A connector integrates with 12 different third-party sources, each with its own OAuth2 app registration and its own rate limit. Credentials for all 12 live in a secrets manager, tagged by source, with the connector's service identity granted read access only to that tag group. A background refresh job checks each source's token expiry every 5 minutes and refreshes any token with under 10 minutes of remaining validity, well before any extraction job would hit it expired. When source #7's refresh token is revoked (an admin at the third-party company deauthorized the app), the refresh job's next attempt fails distinctly, raising a "source 7 needs re-authorization" alert rather than the generic "sync failed" alert every other source's transient hiccup produces, so the on-call engineer knows immediately this needs a human to click through an OAuth consent screen again, not a routine retry.
Trade-offs & pitfalls
- Reactive-only token refresh (refresh on the first 401) is simpler to implement but risks losing an in-flight extraction to a mid-batch expiry, especially for a slow, long-running pull; proactive refresh is worth the extra complexity for anything beyond a trivial connector.
- Logging refresh and retry events "for audit" is only useful if the logs are actually structured and queryable; a wall of unstructured text log lines does not satisfy a real security audit request.
- A single global rate limiter across all sources under-utilizes fast sources and still risks exceeding a slow source's limit; per-source limiting is more code but is the only version that is actually correct.
- Storing a refresh token is itself a long-lived secret with real blast radius if leaked; rotate the underlying OAuth application's credentials periodically even if no incident has occurred, not only in response to one.
Explain pull-based and push-based data ingestion models. For each, give concrete examples (polling a REST API or periodic file fetch versus webhooks or event streams), and compare latency, throughput, operational complexity, load on the source, error and retry behavior, and typical failure modes in production.
Sample Answer
Direct answer
Pull is you initiating contact with the source on your own schedule, for example polling a REST API or fetching a file drop; push is the source initiating contact with you, for example a webhook call or a message it publishes to a stream you subscribe to. Pull gives you full control over pacing and load on the source, at the cost of built-in latency between when something happens and when you notice. Push gives you near-real-time delivery, at the cost of needing to be reliably available to receive it and coordinate with whatever retry behavior the source uses when you are not.
Structured elaboration
Pull
- Concrete examples: polling a REST endpoint every N minutes, fetching a nightly file drop via SFTP (SSH File Transfer Protocol) or from S3, running a scheduled SQL query against a source database.
- Latency: bounded below by your polling interval; a change occurring right after a poll will not be seen until the next one.
- Throughput and source load: you control the request rate directly, which is good for respecting a source's capacity, but a poorly tuned interval can either waste calls when nothing changed or lag badly when a lot changed.
- Operational complexity: you own scheduling, checkpoint tracking, and retry logic; the source does not need to know or care about you.
- Failure modes: a missed poll (your job did not run) simply gets caught on the next poll if your extraction is incremental; the risk is a silent scheduler failure going unnoticed for a while.
Push
- Concrete examples: an inbound webhook call from a payment processor, a message a source publishes to a queue or event stream you consume.
- Latency: near-real-time, since the source notifies you the moment something happens rather than you having to ask.
- Throughput and source load: the source decides the rate, which can spike unpredictably; you need to be able to absorb bursts without falling over.
- Operational complexity: you must run a reliably-available receiver (an endpoint or a consumer), and you inherit whatever the source's own retry and ordering guarantees are, or are not.
- Failure modes: if your receiver is down when a push arrives, you depend entirely on the source retrying it; some sources retry aggressively, some drop the event, and a few offer no redelivery at all.
How to choose
- Freshness requirement: sub-minute or real-time needs generally rule out pure polling.
- Source support: you cannot choose push if the source does not offer it; not every system has webhooks or a stream to subscribe to.
- Control versus availability: pull lets you throttle yourself to protect a fragile source; push demands your receiver be highly available, since you cannot control when the source sends.
- Operational maturity: a small team with no on-call receiver infrastructure may be better served starting with pull, even at some freshness cost, and moving specific sources to push as reliability matures.
Worked example
A team gathering training data for a model has three needs: (a) a large historical backfill of past user actions, (b) online feature updates that must reflect a user's most recent action within seconds, and (c) periodic collection of new human feedback labels. For (a), pull is the only sensible choice: there is no "event" to push, it is a bulk historical extraction, typically against an API or a warehouse export. For (b), push is close to mandatory, since seconds-level freshness is well below what any reasonable polling interval could deliver without hammering the source. For (c), pull on a modest schedule (hourly or daily) is usually sufficient, since new labels do not need to reach the training pipeline instantly, and a scheduled pull is far simpler to operate than standing up a webhook receiver just for this.
Trade-offs & pitfalls
- A common mistake is polling far too aggressively "to reduce latency," which just moves the bottleneck onto the source's rate limits without meaningfully improving freshness once you are polling faster than data actually changes.
- Push without idempotent handling on your side is a duplicate-processing incident waiting to happen, since almost every push-based source will retry a delivery it believes may have failed, even when you actually received and processed it.
- Do not assume push is strictly better because it sounds more modern; a source with unreliable delivery and no replay mechanism can lose data silently in a way a well-designed poll with checkpointing cannot.
- Micro-batching (short, frequent pulls, seconds to low minutes) is a real middle ground worth naming explicitly: it gets you most of push's freshness without needing a highly-available receiver.
You need to backfill two years of historical, paginated data for many accounts from a third-party API that is capped at 10 requests per second, while a live incremental sync keeps running against the same API. Describe your parallelization strategy, how you checkpoint so the backfill can resume, how you coordinate the rate limit across workers, and how you guarantee the result is eventually consistent without duplicating records.
Sample Answer
Direct answer
The core tension is that the backfill and the live sync are both consuming the same 10 requests/second budget from the same API, so the design has to explicitly partition that budget rather than let the two compete unpredictably, while making the backfill itself resumable and idempotent so a multi-day job surviving several restarts still converges on a correct, non-duplicated result.
Structured elaboration
Partitioning the rate-limit budget
- Reserve a fixed share of the 10 req/s for live sync (enough to keep it comfortably within its freshness service-level agreement (SLA)) and let the backfill use the remainder, rather than a naive "whoever gets there first" free-for-all between the two workloads.
- A shared, centralized rate limiter (not one limiter per worker) is required here: independent per-worker limiters cannot see each other's usage and will collectively exceed the source's real cap.
Parallelization strategy
- Parallelize across accounts, not within a single account's page sequence, since pages within one account's history typically must be walked in order for correct checkpointing, while different accounts are fully independent of each other.
- Size the worker pool to the rate-limit budget available to the backfill, not to raw compute capacity; more workers than the rate limit can support just means more of them sitting idle waiting for a token.
Checkpointing for resumability
- Checkpoint per account, independently: persist each account's furthest-completed page or cursor so a restart resumes only the accounts still in progress, not the ones already fully backfilled.
- Persist checkpoints frequently enough (after every page, not just at the end of an account) that a crash loses at most one page of already-fetched-but-uncommitted work per in-progress account.
Guaranteeing eventual consistency without duplicates
- Every record written by either the backfill or the live sync goes through the same idempotent upsert path, keyed on the record's own stable ID, so it does not matter which of the two processes writes a given record first or whether both happen to write it.
- Define a clear ordering rule for the rare case where both processes touch the same record concurrently (for example, "whichever write carries the later
updated_atwins"), so the result is deterministic rather than a race.
Worked example
2,000 accounts need two years of history backfilled, at a shared 10 req/s budget, with 3 req/s reserved for live sync, leaving 7 req/s for the backfill. A pool of 7 backfill workers, each holding one token from a centralized rate limiter, is assigned accounts from a queue; each worker walks one account's full page history, checkpointing its cursor after every page, before picking up the next account from the queue. If the process crashes after 1,200 of the 2,000 accounts are fully done and a 1,201st is half-complete, a restart re-reads the checkpoint table, skips the 1,200 complete accounts entirely, resumes the 1,201st from its last saved cursor rather than page 1, and continues the remaining 799 from scratch. Throughout, both the backfill and the live sync write through the same upsert-by-record-ID path, so if live sync happens to ingest a very recent change to an account the backfill has not yet reached, the backfill's later arrival at that same record, carrying older data, does not overwrite the newer one, because the upsert conflict rule explicitly prefers the later updated_at.
Trade-offs & pitfalls
- Splitting the rate-limit budget statically (a fixed 3/7 split) is simple but can starve one workload if the actual demand shifts, for example if live sync traffic spikes; a smarter allocator that lets the backfill temporarily borrow idle live-sync capacity is more efficient but meaningfully more complex to build and reason about correctness for.
- Checkpointing at the account level but not the page level within an account means a crash can lose up to one full account's progress, not just one page; the finer-grained the checkpoint, the less repeated work a crash costs, at the price of more frequent writes to the checkpoint store.
- The "later
updated_atwins" conflict rule assumes the source's own timestamps are trustworthy and consistent across accounts and over the two-year backfill window; if that assumption does not hold for some historical data, you need a different, explicit tie-breaking rule rather than silently trusting a timestamp that might be wrong. - A single centralized rate limiter is also a single point of contention; at very high worker counts it can itself become a bottleneck, though at 7 concurrent workers against a 7 req/s budget this is not yet a real concern.
Unlock Full Question Bank
Get access to all 18 Data Ingestion and Source System Integration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.