Rate Limiting, Throttling and Quota Management Questions
Protecting API capacity and enforcing fair use: rate-limiting algorithms (token bucket, leaky bucket, fixed/sliding window), per-client quotas, throttling responses (429 semantics, Retry-After), and tiered plan enforcement. Covers where to enforce limits (gateway vs. service), distributed counters, and graceful degradation under load.
Implement an in-memory per-client rate limiter in a language of your choice (suggested: Python or Go). The limiter should support the token bucket algorithm with configurable refill rate and capacity, be concurrency-safe, and provide an API allow(client_id) -> bool. Discuss memory usage for large numbers of clients and strategies to evict inactive clients.
Compare rate limiting algorithms: token bucket, leaky bucket, fixed window, sliding window log, and sliding-window counter. Explain which algorithms work best for bursty traffic, how to approximate sliding windows in distributed systems, and how to apply rate limits per user, per API key, and globally.
Design a distributed rate limiter for a login API that must prevent brute-force attacks: per-IP limit 20 req/min, per-account limit 10 req/min, plus a global emergency throttle. Describe algorithms, storage choices (in-memory, Redis, CRDTs), how to enforce limits across regions, and strategies to avoid false positives for users behind NAT or large proxies.
That is every published Rate Limiting, Throttling and Quota Management question for Systems Engineer so far. Browse the other topics in this category, or practice this one interactively.