What are different rate limiting algorithms?

Image

Rate limiting is a crucial technique in controlling the amount of traffic a server receives within a specified time frame. It's used to prevent overuse of resources, improve server reliability, and ensure fair usage among users. Rate limiting is common in API management to prevent abuse and to manage traffic effectively.

Different Rate Limiting Algorithms:

1. Fixed Window Counter

  • Description: Divides time into fixed windows and counts the number of requests in each window.
  • Example: If the limit is 100 requests per hour, and a user makes 100 requests in the first half-hour, they will be blocked for the remaining half-hour, even if the server is underutilized during that time.

2. Sliding Log

  • Description: Keeps a time-stamped log of requests. It checks whether adding a new request would exceed the rate limit, considering the time frame.
  • Example: If the limit is 100 requests per hour, each incoming request is checked against the log of requests in the past hour. Older entries are discarded.

3. Sliding Window Counter

  • Description: A hybrid of the fixed window and the sliding log, offering a balance between efficiency and precision. It combines the fixed window's simplicity and the sliding log's accuracy.
  • Example: If the limit is 100 requests per hour, the server counts requests in the current window and a fraction of the requests from the previous window, based on the time elapsed.

4. Token Bucket

  • Description: Uses tokens to control traffic flow. Tokens are added to a bucket at a regular rate and requests consume tokens. If the bucket runs out of tokens, new requests are denied.
  • Example: A bucket can hold 10 tokens and 1 token is added every 10 seconds. A request needs 1 token to pass. If there's a sudden burst of 15 requests, only 10 can go through, and subsequent requests must wait for new tokens.

5. Leaky Bucket

  • Description: Requests are added to a queue (bucket) and processed at a fixed rate to smooth out burst traffic.
  • Example: If the bucket size is 10 and the rate is 1 request per second, and a burst of 20 requests comes in, the first 10 are queued and processed at 1 per second, while the rest are either queued (if the bucket can hold them) or discarded.

Application of Rate Limiting

  • APIs and Web Services: To control traffic and prevent abuse.
  • Network Traffic: To control data flow in networks.
  • Application Servers: To prevent overload and ensure fair usage.

In implementing rate limiting, it's crucial to choose an algorithm that aligns with the system's needs, balancing between fairness, efficiency, and resource utilization.

Distributed Rate Limiting

In a microservices or multi-node deployment, the counter cannot live in one process. The usual fix is a shared store such as Redis that holds the request counts or the token state for every node.

  • Pros: works when requests arrive through many nodes, and every node sees the same limit.
  • Cons: needs a fast and reliable central store, which becomes a bottleneck and a single point of failure if it is not planned for.

Check out how to design Rate Limiter.

Learn about system design concepts in Grokking the System Design Interview course.

Rate limits are also part of the API contract. Grokking Modern API Design Interview covers the headers and error responses that tell a client its limit, and how to limit expensive compute such as model inference.

TAGS
System Design Fundamentals
Scalability
Rate Limiting
CONTRIBUTOR
Arslan Ahmad
Arslan Ahmad
ex-FAANG engineering manager and author or Grokking series.

GET YOUR FREE

Coding Questions Catalog

Design Gurus Newsletter - Latest from our Blog
Boost your coding skills with our essential coding questions catalog.
Take a step towards a better tech career now!
Explore Answers
What to Expect in the Palantir System Design Interview
Palantir's design evaluation centers on the Decomposition round: a vague real-world problem you must turn into buildable structure. Strategy, reported prompts, and the data-heavy SD variant.
Explain gRPC Streaming vs Unary Calls.
The four gRPC call types, what streaming changes for load balancing and cleanup, and when each one is the right choice.
What is Single-Tenant vs Multi-Tenant Systems?
What to Expect in the Cohere System Design Interview
Cohere design rounds center on enterprise AI serving: multi-tenant models at low latency, RAG with measured quality, and deployment into customer environments. Themes and preparation.
System design resources recommended by top tech companies
Discover the system design resources that FAANG engineers actually use and recommend. Covers engineering blogs, open-source papers, courses, books, and internal training materials made public.
Forward Proxy vs Reverse Proxy: What Is the Difference?
A forward proxy acts for the clients behind it; a reverse proxy acts for the servers behind it, and every other difference follows from that.
Related Courses
New
Grokking the AI System Design Interview course cover
Grokking the AI System Design Interview
Learn to design AI systems the way interviewers expect: classic ML products, LLM and RAG architectures, and agentic systems, all through the lens of the system design interview.
4.6
(3,192 learners)
Discounted price for Your Region

$99

Grokking the Coding Interview: Patterns for Coding Questions course cover
Grokking the Coding Interview: Patterns for Coding Questions
The 24 essential patterns behind every coding interview question. Available in Java, Python, JavaScript, C++, C#, and Go. The most comprehensive coding interview course with 543 lessons. A smarter alternative to grinding LeetCode.
4.6
Discounted price for Your Region

$197

Grokking Modern AI Fundamentals course cover
Grokking Modern AI Fundamentals
Master the fundamentals of AI today to lead the tech revolution of tomorrow.
4.1
Discounted price for Your Region

$72

Design Gurus logo
One-Stop Portal For Tech Interviews.
Copyright © 2026 Design Gurus, LLC. All rights reserved.