We Were Rate-Limiting Our Own Services
Our API's per-address rate limit treated our own services like one busy person, and in production they started getting 429s. How we found it and fixed it.
Mythex Team · · 4 min read
Mythex's API allows 300 requests a minute from one network address. That limit exists to protect us from a single person or script hammering the API. In production, it started refusing our own services instead — the part of our edge that serves published apps and the service that runs the AI agent — because each of them reaches the API from only a handful of addresses. Requests that carry our internal credential are now exempt, and everyone else is limited exactly as before.
The symptom
The API began answering some of our own internal calls with 429 Too Many Requests. Three kinds of call were hit, and each one matters to someone using Mythex:
- Looking up which app lives on a custom domain. When a visitor opens a published app on its own domain, our edge asks the API where that site is. A 429 there is a visitor who can't reach a site.
- Saving a running agent turn. While the agent works, its progress is saved as it goes, so a reply survives a refresh or a dropped connection. A 429 there is progress that didn't get saved.
- Checking access to a preview. The preview of an app you're building is gated, and the gate asks the API. A 429 there is a preview that won't open.
None of these were people calling the API. All of them were us.
Why it happened
A per-address rate limit makes one assumption: one address is roughly one caller. For people, that's close enough. For our own infrastructure, it's badly wrong.
- The edge service that serves published apps and previews reaches the API from a few shared outgoing addresses. Every visitor to every published app, and every preview check, arrives looking like it came from one of those few.
- The orchestrator — the service that runs agent turns — reaches the API from one container address. Every running turn, for every user, shares it.
So the budget meant for one person was being split across all the traffic of an entire service. Those services used it up between them, and the limiter did what it was told: it refused them.
The fix
Our own services already prove who they are. Every internal call carries an internal token, and every route under the API's internal section checks that token before it does anything.
The rate limiter now uses the same check. A request whose token matches the internal one is on an allow list and isn't counted against the per-address budget. Everything else is counted exactly as it was: 300 requests a minute per address.
Two details kept this from becoming a hole:
- A wrong token gets no pass. The allow list only matches the real internal token. A request with a made-up token is counted and throttled like any other, and the internal routes still reject it with
401 Unauthorized. - The limit itself didn't change. We didn't raise the number to make room. Raising it would have hidden the problem until traffic grew again, and it would have weakened the protection for the case the limit exists for.
We moved the limiter's settings into one small function and wrote tests against it. One sends 20 requests with the internal token to a limiter set to allow 5, and expects all 20 to succeed. Another sends plain requests from one address and expects the first five to succeed and the ones after to be refused. A third sends requests with a forged token and expects them to be refused the same way.
What we took from it
- Limits carry assumptions. "One address, one caller" is built into a per-address limit. It stops being true the moment your own services sit behind shared addresses — and many do: load balancers, edge networks, containers, office networks.
- Tell your own traffic apart by identity, not by address. Addresses are how traffic arrives, not who sent it. We already had a credential that proves a call is ours; the limiter just wasn't looking at it.
- Don't fix it by raising the number. A bigger limit would have made the 429s go away for a while, protected less, and failed again later.
- Look at who's getting throttled. A 429 that lands on your own service is a bug report about your limits, not a sign the limiter is working.
What it means if you build on Mythex
Nothing to change on your side. Published apps on custom domains, agent turns and previews no longer depend on how busy the rest of Mythex is. Your app's own traffic is your app's business: the limit described here is on our API, not on the apps you publish.
If you run a backend of your own, the same trap applies to any rate limit you add. Our guides on what an API is and the security checklist for AI-built apps are a good place to start. And for another case of our own infrastructure getting in its own way, see how we got deploys down to zero 502 errors.