Product Analytics Instrumentation and Event Tracking Questions
Instrumenting products to collect behavioral data: event taxonomy/tracking plans, client and server-side collection, attribution implementation, and telemetry for web, mobile, and games (including crash reporting). Covers designing clean, analyzable event schemas and the collection infrastructure behind them. The data-collection foundation for product analytics.
Design a cross-platform analytics and instrumentation pipeline that aggregates events from iOS and Android, supports performance monitoring and crash grouping, respects user privacy and GDPR requests, and enables funnel and cohort analysis. Specify SDK choices or alternatives, a recommended event schema and versioning approach, sampling strategies, data retention policies and how to guarantee no PII is recorded.
Sample Answer
High-level approach
- Use a lightweight cross-platform SDK layer in-app that forwards events to a telemetry backend (e.g., Segment or an in-house collector). Backends: Firebase Analytics + Crashlytics (mobile-first), Sentry/Datadog for performance & crash grouping, or OpenTelemetry + Snowflake/BigQuery for custom pipelines. Keep SDK usage optional: primary analytics via Segment (router) -> destinations (BigQuery, Sentry).
SDK choices / alternatives
- Primary: Segment (JS/React Native/Swift/Kotlin) or RudderStack -> BigQuery for funnels/cohorts.
- Crash/perf: Sentry or Firebase Crashlytics (both support symbolication, grouping).
- Observability: OpenTelemetry + Datadog for traces/metrics.
- Reason: Segment centralizes consent/sampling, Sentry provides crash grouping and perf traces.
Recommended event schema & versioning
- Always include schema_version, event_type, timestamp (ISO8601), platform, sdk_version, session_id.
- Minimal required properties: pseudonymous_user_id, anon_id, event_name, event_props (object), revenue (optional), device_os_version, app_version.
- Versioning: increment schema_version on breaking changes; keep backward compatibility; include deprecated_fields array if needed.
Example event:
{
"schema_version": 2,
"event_type": "purchase_completed",
"timestamp": "2026-02-01T12:34:56Z",
"platform": "iOS",
"app_version": "1.4.0",
"pseudonymous_user_id": "user_XXXX_hash",
"session_id": "sess_abc",
"event_props": {"product_id":"sku_123","price":9.99}
}
Sampling strategies
- Client-side: low-frequency events sampled (e.g., 1-5%) with deterministic hashing on anon_id to keep cohort stability.
- Server-side: sample high-volume event types in ingestion; always keep 100% for crashes, conversion events, and performance spans above thresholds.
- Adaptive sampling: increase sample rate for new releases or anomalies.
Data retention & storage
- Raw events: keep 90 days in hot store (BigQuery/Databricks), then aggregate to weekly/monthly summaries kept 2+ years.
- Crash payloads: keep full symbolicated crash data 1 year, aggregated crash fingerprints 3+ years.
- Audit logs and deletion requests: retain 1 year for compliance.
- Encrypt data at rest and in transit.
GDPR & privacy
- Consent-first: require explicit consent toggle before any tracking; store consent state in SDK and server.
- Right-to-be-forgotten: provide API to delete/pseudonymize user data (delete event rows or replace pseudonymous_user_id with tombstone token).
- Data minimization: default to anonymous tracking; only collect user_id after opt-in.
- Consent propagation: include consent metadata with every event and drop events lacking consent for sensitive categories.
Guaranteeing no PII
- Prohibit capturing free-text user input; implement client-side allowlist of fields.
- Client SDK enforces schema and runs scrubbers:
- Regex scrubbers for emails, phone numbers, credit-card patterns - redact client-side.
- Field allowlist: only allow specific keys (product_id, category, price). Any other keys are rejected.
- Hashing: if a stable identifier is required, client-side HMAC-SHA256 with app-specific salt before sending (no raw emails/usernames).
- Server-side validation: reject/blackhole events containing PII patterns; log and alert.
- Periodic audits and automated tests to ensure no PII leakage.
Funnel & cohort support
- Use event_name + consistent properties (pseudonymous_user_id, session_id, timestamp) to construct funnels.
- Cohorts built from historical event_props and attributes in BigQuery; keep cohort membership snapshots daily.
- Maintain deterministic anon_id hashing to keep cohort stability across sampling.
Trade-offs / reasoning
- Centralized router (Segment) simplifies consent, sampling and destination control.
- Sentry/Crashlytics specialized crash grouping vs custom pipeline: choose managed for faster delivery; custom if you need deep analytics integration.
- Client-side scrubbing reduces GDPR risk; server-side checks are last defense.
This design balances mobile constraints (battery, network), GDPR requirements, accurate funnels/cohorts, and reliable crash/perf insights.
Explain how Data Engineers, Product Managers, and Data Scientists should collaborate to define instrumentation and event schemas for a new product launch. Provide a checklist (event names, payload required fields, identifiers, timestamp format, failure modes, test plans) that must be agreed before release.
Sample Answer
Collaboration approach (how we work together)
- Kickoff: PM defines product goals, key experiments/metrics (MAU, conversion, retention) and success criteria. Data Scientist lists required signals and derived metrics. Data Engineer proposes feasible capture points, schema standards, and retention/throughput constraints.
- Joint design session: map user flows to events, agree ownership, enforce naming and field contracts, and add observability hooks (sampling, throttles).
- Sign-off & QA: PM signs behavioral intent, DS signs analytical sufficiency, DE signs implementation and operational constraints.
- Iteration: instrument in staging, run validation tests, roll out progressively (feature flags, percentage rollout).
Pre-release checklist (must be agreed and documented)
- Event names (convention: snake_case, prefix by domain)
- e.g., product_view, signup_started, purchase_completed
- Event payload required fields
- event_type (string)
- user_id (stable, anonymized if needed)
- session_id (UUID)
- product_id / feature_id (canonical IDs)
- event_version (int)
- event_props (object/map with documented typed fields)
- Identifiers & identity resolution
- user_id (primary), anon_id (cookie/device), account_id (if B2B)
- specify format, hashing/encryption rules, PII handling
- Timestamp format
- event_timestamp ISO8601 UTC with millisecond precision (e.g., 2025-12-06T14:23:12.123Z)
- server_received_at for ingestion latency tracking
- Failure modes & validation
- missing required field -> drop or route to error topic with reason
- malformed payload -> reject and store raw for debugging
- duplicate events -> include event_id (UUID) and idempotency handling
- high-volume spikes -> sampling policy and backpressure behavior
- Schema versioning & governance
- backward-compatible additions allowed; breaking changes require migration plan and version bump
- central schema registry (JSON Schema/Avro/Protobuf)
- Observability & SLAs
- metric dashboards: event volume, schema validation failure rate, ingestion lag
- alert thresholds (e.g., >1% schema errors)
- Test plan (signed by DE & DS)
- unit tests for event serializers/deserializers
- staging end-to-end tests: simulate user flows producing events; validate presence, schema, timestamps
- contract tests against schema registry
- load test for expected peak QPS + 2x
- backfill test for late-arriving events handling
- Release & rollback plan
- feature-flagged rollout (10% → 50% → 100%)
- monitoring checks at each stage and rollback criteria
- Documentation & handover
- canonical event spec with examples, required fields, types, allowed values, and sample payloads stored in central doc repo
As Data Engineer I drive schema enforcement, implement ingestion pipelines, provide test harnesses and dashboards, and own rollback/operational playbooks. Data Scientists validate analytical sufficiency; PM ensures product intent and acceptance criteria. Agreement on the checklist before release prevents costly rework and ensures reliable analytics from day one.
Create a plan for instrumenting a new feature to make sure business KPIs (e.g., average order value, conversion) are captured end-to-end. Include event schema design, sampling strategy, QA checks, and how you'll report results to product and analytics teams.
Sample Answer
Plan overview: instrument the new checkout-promotion feature end-to-end so business KPIs (AOV, conversion) are reliable and auditable. Phases: design, implement, test/QA, rollout with sampling, monitor & report.
- Event schema design (single source of truth)
Provide consistent, minimal events with clear IDs and types. Example canonical events:
{
"event_type": "promotion_applied",
"timestamp": "ISO8601",
"user_id": "anon|<hashed_id>",
"session_id": "<uuid>",
"order_id": "<nullable>",
"product_ids": ["p1","p2"],
"price_cents": 1999,
"discount_cents": 300,
"promotion_id": "PROMO123",
"channel": "web|ios|android",
"metadata": {"ab_variant":"B"}
}
Also send order_created and order_completed events with identical order_id and price/line-item breakdown.
- Sampling strategy
- Default: capture 100% of conversion-critical events (order_created/completed, promotion_applied).
- For high-volume exploratory events (page_views), sample at 5–10% and annotate sampled=true and sampling_rate. Ensure sampled events still include session_id to link to full events.
- QA checks
- Schema validation at ingestion (reject/log malformed events).
- Synthetic test suite: replay deterministic test users through feature to assert sequence: promotion_applied → order_created → order_completed with same order_id.
- Consistency checks: price math validation (sum(line_items) - discount == order_total).
- Duplicate detection: same event signature within short window flagged.
- Instrumentation smoke dashboards: event counts per minute, drop-off funnel (applied→created→completed).
- Canary rollout: enable for 1% of users, verify metrics before scaling.
- Reporting to product & analytics
- Provide a rollout brief with event dictionary, sampling notes, and known caveats.
- Build 3 dashboards: funnel conversion (applied→completed), AOV by cohort (variant/channel), and anomaly charts with alerts (5% change or p < 0.01).
- Weekly sync for first month with product and analytics to review metrics, edge cases, and iterate.
- Deliver reproducible SQL queries and notebooks that join events on order_id/session_id so analysts can validate.
Rationale: explicit schema + 100% capture for critical events ensures KPI accuracy; sampling saves cost for noisy events; QA and canary reduce risk; clear reporting and reproducible artifacts make results actionable.
A new regulation requires deletion of user data upon request within 48 hours. Describe how you would design retention and deletion pipelines to meet this requirement for both online and archived event stores, and how you would validate deletion completion for audits.
Sample Answer
Requirements & constraints:
- Functional: honor user deletion requests within 48 hours across online stores (user profiles, live event tables, caches) and archived event stores (data lakes, S3, cold backups).
- Non-functional: provable audit logs, minimal service disruption, low false positives, scalable to N deletions/day.
- Regulatory: soft vs hard delete policy clarification; whether downstream derivatives (analytics, ML models) must be purged.
High-level design:
- Request intake & orchestration
- Self-serve portal + API that creates a deletion-job with user ID, scope, timestamp, SLA (48h).
- Authorization & verification workflow (consent checks, legal holds).
- Deletion engine
- Online pipeline: Fan-out orchestrator issues targeted deletes to microservices/databases and caches. Use idempotent delete APIs and CDC to remove events referencing user_id.
- Archive pipeline: Cataloged retention index maps user_id -> data locations (tables, S3 prefixes, partitions). Batch jobs (Spark) run targeted deletes or rewrite partitions to exclude user data. For immutable object stores, use tombstone metadata + lifecycle policy to purge physical objects within retention window.
- Message bus (Kafka) for reliable commands and status updates; each subsystem emits completion events.
- Auditing & verification
- Append-only audit log with job lifecycle, signed timestamps, operator IDs.
- Validation service: sampling + deterministic scans (hash-based) to assert absence in indexes and sample objects. Produce a signed compliance report per job.
- Reconciliation: periodic end-to-end job scanner that cross-checks catalog, DB indices, and object storage for residual PII.
Key considerations & trade-offs:
- Immediate consistency vs eventual: online deletes should be synchronous for profiles; analytics/archives can be eventual but must complete within 48h.
- Cost: rewrites of large partitions are expensive - prefer targeted object-level deletes when possible; maintain mapping index upfront to avoid full-scan deletions.
- Legal hold override path with strict logging.
Operational plan & metrics:
- KPIs: % deletions completed within 48h, audit verification pass rate, mean time to detect residuals.
- Roadmap phases: 1) intake + online deletes, 2) archival indexing + targeted archive deletes, 3) audit/validation automation, 4) scale & hardening.
This approach balances compliance, scalability, clear ownership, and auditable evidence for regulators.
Explain the difference between sampling at ingestion vs. sampling at query/analytics time. For a high-volume telemetry stream, when would you choose each approach and what are the implications for OLAP accuracy and cost?
Sample Answer
Summary:
Sampling at ingestion drops or reduces data when it first arrives; sampling at query/analytics time retains raw/near-raw data and samples only when answering a query. Each approach trades cost, flexibility, and analytic accuracy differently.
Sampling at ingestion
- When to choose: very high-volume telemetry with tight storage/ingest budget, stable/known KPIs, and mostly pre-defined aggregate dashboards (e.g., simple counts, rates).
- Pros: big savings in storage, indexing, and long-term compute; simpler pipeline; lower downstream costs.
- Cons: irreversible loss of fidelity - you can’t run new analyses later that require discarded data; introduces sampling bias if not carefully designed.
- OLAP implications: lower cost but reduced accuracy and higher variance for low-frequency events; need to use appropriate weighting/estimation and document the sampling scheme for correct interpretation.
Sampling at query/analytics time
- When to choose: exploratory analytics, anomaly detection, compliance/audit needs, or when product teams expect evolving questions about the data.
- Pros: maximum analytical flexibility; unbiased results if sampling is done correctly at query time; supports ad-hoc, retrospective analyses.
- Cons: higher storage and ingestion costs; larger query compute (may require more aggressive query sampling/approximation techniques).
- OLAP implications: better accuracy/fidelity and ability to compute confidence intervals per query, but cost scales with retention and query complexity.
Practical patterns / hybrid options
- Store full data for a small fraction (100% for 1% of users or errors) + ingestion-sample the rest - preserves rare-event analysis while cutting cost.
- Use stratified/deterministic sampling keyed by user/session/error severity to reduce bias.
- Retain raw for a short hot window (e.g., 7 days) then sample older data.
- Attach sampling metadata (sample rate, key strata) at ingestion so analysts can weight results correctly.
Recommendation (product POV)
- If customers require deep, evolving analytics or regulatory auditability, favor query-time sampling (invest in cost controls: tiered storage, cold paths).
- If cost is a dominant constraint and analytics needs are fixed and well-understood, use ingestion sampling with careful design, clear documentation, and hybrid safeguards to protect rare-event visibility.
Unlock Full Question Bank
Get access to all 25 Product Analytics Instrumentation and Event Tracking interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.