Guides / Concepts explained

What Is Rate Limiting? A Plain-English Guide for App Builders

Rate limiting caps how many requests someone can make in a set time. Why it protects your app and your bills, how 429 errors work, and what to ask your AI.

Mythex Team · 2026-09-29 · 6 min read

Rate limiting is a rule that caps how many requests someone can make to a system in a set amount of time — for example, 5 login attempts per minute, or 100 API calls per hour. Once the limit is reached, further requests are refused (usually with the error code 429 Too Many Requests) until the time window resets. It protects apps from abuse, keeps one heavy user from slowing things down for everyone, and stops runaway costs.

You meet rate limiting from two sides: services you call limit you, and your own app should limit its users.

Why rate limiting matters when you build with AI

AI-built apps often connect to paid services — AI models, email, SMS, maps — and each call costs money. That makes rate limiting both a security feature and a budget feature.

  • Cost control. If your app has an "ask AI" box and no limit, a bot can send it thousands of requests overnight. You pay for every one.
  • Security. Without a limit on the login form, an attacker can try thousands of passwords against an account. A limit of a few attempts per minute makes that impractical.
  • Staying online. One misbehaving script — sometimes your own code stuck in a loop — can overload your database and take the app down.
  • Respecting other services' limits. When your app calls an outside API too often, you get the 429 errors, and features break for your users.

AI builders will build what you ask for. If you don't mention limits, they're often left out.

An everyday analogy

Think of a coffee shop's free refill policy: "One free refill per customer per visit."

  • It doesn't stop anyone from enjoying coffee.
  • It stops one person from standing at the counter all day refilling a flask for the whole office.
  • When someone asks for a fourth refill, the barista politely says "not right now" — that's the 429.

A good rate limit is invisible to normal customers and only noticed by someone doing something unusual.

How rate limiting works

A rate limiter needs three things:

PartQuestion it answersExamples
WhoWhose requests are being counted?Per user account, per IP address, per API key
How manyWhat's the limit?5, 60, 1,000 requests
Per what timeOver what window?Per second, minute, hour, day

When a request arrives, the app checks the count for that "who." Under the limit: the request goes through and the count goes up. Over the limit: the request is refused.

There are a few common ways to count:

MethodHow it worksTrade-off
Fixed windowCount resets on the clock — e.g. every minute on the minuteSimple, but allows a burst at the edge of two windows
Sliding windowLooks at the last 60 seconds from nowSmoother, slightly more work to compute
Token bucketEach user has a bucket that refills at a steady rate; each request takes a tokenAllows short bursts while holding the average steady

The counts need to be stored somewhere fast and shared. On a single small server, memory works. Once an app runs on several machines, the counter usually lives in a shared store like Redis or the database, otherwise each machine counts separately and the real limit becomes several times higher.

What a rate-limited response looks like

When a request is over the limit, a well-behaved API responds with:

  • Status code 429 Too Many Requests.
  • Often a Retry-After header, saying how many seconds to wait.
  • Sometimes headers showing your limit and how many requests you have left in this window (names vary by provider).

Your app should respond to a 429 by waiting and trying again later — ideally with exponential backoff, which means waiting a little longer after each failed attempt (1 second, then 2, then 4) instead of hammering the service.

A worked example: an AI writing assistant

Say you've built an app where users paste text and get an AI rewrite.

Without limits: someone writes a script that sends 20,000 requests overnight. Your AI provider bill jumps, and your account hits the provider's own rate limit, so real users see errors all morning.

With sensible limits:

  1. Per signed-in user: 30 rewrites per hour on the free plan, 300 on paid. Over the limit, the user sees: "You've reached this hour's limit. Try again in 12 minutes, or upgrade."
  2. Per IP address on the sign-up form: 5 sign-ups per hour, to stop bots creating thousands of free accounts to get around the per-user limit.
  3. A spending cap set in your AI provider's dashboard, as a last line of defence if everything else fails.
  4. Handling the provider's limits: if the AI provider returns a 429, your server waits and retries once or twice, then shows a friendly "busy, please try again" message.

The limits are checked on the server. A limit that only exists in the browser — for example, disabling the button after 30 clicks — can be bypassed by anyone who calls your API directly.

Key terms

TermMeaning
Rate limitThe maximum number of requests allowed in a time window.
QuotaA larger allowance, often per day or month, sometimes tied to a paid plan.
429The HTTP status code for "too many requests."
Retry-AfterA response header saying how long to wait before trying again.
ThrottlingSlowing down or refusing requests over a limit. Often used as a synonym.
Exponential backoffWaiting progressively longer between retries.
BurstA short spike of many requests at once.
Brute-force attackTrying huge numbers of passwords or codes until one works.

Common misconceptions

  • "Rate limiting is only for big apps." Small apps are targets too, because bots scan the whole internet for open forms and AI endpoints. The cost of abuse doesn't depend on your size.
  • "Hiding the button is enough." Browser-side limits are a courtesy, not protection. Real limits must be enforced on the server.
  • "Per-IP limits are always fair." Many people can share one IP address — an office, a school, a mobile network. Per-user limits are fairer once people are signed in.
  • "A 429 means the service is broken." It means you're calling it too often. Check your code for loops, and cache results that don't change often — see what caching is.
  • "Rate limiting replaces a spending cap." Use both. Limits in your app can have bugs; the provider's spending cap is your backstop.

What to ask your AI builder

  • "Add server-side rate limiting to login, sign-up, password reset and the AI endpoint. Suggest sensible limits."
  • "Limit each signed-in user to [N] AI requests per hour, and show a clear message with the wait time when they hit it."
  • "When an outside API returns 429, retry with exponential backoff up to 3 times, then show a friendly error."
  • "Where are the rate-limit counters stored? Will they still work if the app runs on more than one machine?"
  • "Log whenever a user or IP hits a limit, so I can spot abuse."
  • "Which outside services does this app call, and what are their rate limits?"

Rate limiting in Mythex

Mythex's docs are direct about this for AI features: rate limits, safety and cost controls in apps you build are yours to design — see Build an AI app. Ask the agent to add limits in the same message where you ask for the feature. On Pro, a published service can run on up to three machines when traffic is heavy (see Hosting limits), so ask for counters that live in the database rather than in one server's memory.

Related reading: what an API is, how to use LLM APIs and the security checklist for AI-built apps.

Questions

What is rate limiting in simple terms?

Rate limiting is a rule that caps how many requests a user, device or app can make in a given period — for example, 5 login attempts per minute or 100 API calls per hour. Requests over the limit are refused until the window resets.

What does a 429 Too Many Requests error mean?

HTTP status 429 means you sent too many requests in too short a time and hit a rate limit. The response often includes a Retry-After header saying how long to wait. The fix is to slow down, retry after a pause, or reduce how often your app calls the service.

Why does my AI-powered app need rate limiting?

Every AI request costs you money in tokens. Without a limit, one user or a bot can send thousands of requests and run up a large bill or use up your provider quota, taking the feature down for everyone. Per-user limits keep costs predictable.

Is rate limiting the same as throttling?

They're closely related and often used interchangeably. Rate limiting usually means rejecting requests over a limit, while throttling can also mean slowing them down or queuing them instead of refusing them outright.

Does rate limiting stop DDoS attacks?

It helps against small floods and abuse from individual users or IP addresses, but a large distributed attack comes from many addresses at once. Protection against those usually happens at the network or CDN level, before traffic reaches your app.

Keep reading

  • Frontend vs Backend: What's the Difference? — The frontend is what users see in the browser; the backend runs on a server and handles data, logic and security. How the two fit together, with an example.
  • How Domains and DNS Work: A Guide for Non-Developers — How domain names and DNS connect example.com to your app: registrars, nameservers, A, CNAME, MX and TXT records, propagation, and connecting a custom domain.
  • How to Use LLM APIs: Tokens, Costs, Keys and Your First AI Feature — What an LLM API is, how tokens, context windows and per-token pricing work, how to keep your API key safe, and how to add a first AI feature to your app.
  • Native Apps vs Progressive Web Apps: Which Do You Need? — Native apps vs progressive web apps (PWAs): what each can do, iPhone limits as of September 2026, costs, and how to choose for your first version.
  • REST vs GraphQL: What's the Difference and Which Should You Use? — REST and GraphQL are two ways to design an API. How each works, with examples, the real trade-offs, and which one makes sense for an app you build with AI.
  • Security Checklist for AI-Built Apps: What to Check Before Launch — A practical security checklist for AI-built apps: login and permissions, secrets, input validation, per-user data, rate limits, dependencies and backups.

Start building free · Templates · Docs