Overview
The rate limiting example demonstrates two types of rate limiting:- Message Rate Limiting: Prevents users from sending messages too frequently
- Token Usage Rate Limiting: Controls AI model token consumption over time
Running the Example
Rate Limiting Strategy
Below we’ll go through each configuration. You can also see the full example implementation in rateLimiting.ts.1. Fixed Window Rate Limiting for Messages
- Allows 1 message every 5 seconds per user.
- Prevents spam and rapid-fire requests.
- Allows up to a 2 message burst to be sent within 5 seconds via
capacity, if they had usage leftover from the previous 5 seconds.
- Allows 1000 messages per minute globally, to stay under the API limit.
- As a token bucket, it will continuously accrue tokens at the rate of 1000 tokens per minute until it caps out at 1000. All available tokens can be used in quick succession.
2. Token Bucket Rate Limiting for Token Usage
- Allows 1000 tokens per minute per user (a userId is provided as the key), and 100k tokens per minute globally.
- Provides burst capacity while controlling overall usage. If it hasn’t been used in a while, you can consume all tokens at once. However, you’d then need need to wait for tokens to gradually accrue before making more requests.
- Having a per-user limit is useful to prevent single users from hogging all of the token bandwidth you have available with your LLM provider, while a global limit helps stay under the API limit without throwing an error midway through a potentially long multi-step request.
How It Works
Step 1: Pre-flight Rate Limit Checks
Before processing a question, the system:- Checks if the user can send another message (frequency limit)
- Estimates token usage for the question
- Verifies the user has sufficient token allowance
- Throws an error if either limit would be exceeded
- If the rate limits aren’t exceeded, the LLM request is made.
limit and check is that limit will consume the
tokens immediately, while check will only check if the limit would be
exceeded. We actually mark the tokens as used once the request is complete with
the total usage.
Step 2: Post-generation Usage Tracking
While rate limiting message sending frequency is a good way to prevent many messages being sent in a short period of time, each message could generate a very long response or use a lot of context tokens. For this we also track token usage as its own rate limit. After the AI generates a response, we mark the tokens as used using the total usage. We usereserve: true to allow a (temporary) negative balance, in case
the generation used more tokens than estimated. A “reservation” here means
allocating tokens beyond what is allowed. Typically this is done ahead of time,
to “reserve” capacity for a big request that can be scheduled in advance. In
this case, we’re marking capacity that has already been consumed. This prevents
future requests from starting until the “debt” is paid off.
When using the Agent component, we can do this in the “usageHandler”, which is
called after the AI generates a response.
Client-side Handling
See RateLimiting.tsx for the client-side code. While the client isn’t the final authority on whether a request should be allowed, it can still show a waiting message while the rate limit is being checked, and an error message when the rate limit is exceeded. This prevents the user from making attempts that are likely to fail. It makes use of theuseRateLimit hook to check the rate limits. See the full
Rate Limiting docs here.
bijection/example.ts we expose getRateLimit: