429 Too Many Requests.
To make your app production-ready, use these four patterns to prevent this or handle errors when it happens:
- Client-side caching for hot domains.
- Backoff on 429, honoring the
Retry-Afterheader. - Prefetch to shift slow work ahead of bursts.
- Tier-aware fallbacks when the limit holds.
Rate limits per plan
Rate limits apply per API key, are measured per minute, and are visible on your dashboard. The current tiers:
* Free plan credits are a one-time grant, not a monthly allowance.
Logo Link and Prefetch endpoints do not have any rate limits. Monitors management endpoints (
/v1/monitors/*) use a separate per-organization bucket.Monitors API has its own bucket
Requests to/v1/monitors/* — the endpoints that create, list, update, delete, and inspect monitors and their runs — do not count against the per-plan rate limit in the table above. They draw from a separate per-organization bucket:
This bucket is isolated in both directions: heavy monitor polling never eats into your data-API budget, and heavy brand/web traffic never throttles your monitors dashboard or webhook backfill queries. The 1000/min ceiling is flat across every plan, including Free.
When you exceed the monitors bucket, the response is the same
429 Rate limit exceeded envelope described below, but the X-RateLimit-* headers on monitors responses reflect the monitors bucket (limit 1000) rather than your plan’s data-API limit. Apply the same backoff pattern — reading Retry-After and X-RateLimit-Remaining from the monitors response works the same way.
Monitor runs (the actual scrape work executed on your monitor’s schedule) are separate from the monitors management API. Run frequency and count are controlled by your plan’s monitor limits and the monitor’s own schedule, not this per-minute cap.
Weighted endpoints
Two endpoints fan out to many upstream scrapes per call and count as 10 requests against your per-minute rate limit instead of 1:
Every other endpoint still counts as 1 request. On a Pro plan (300 requests/min), that means at most 30 crawl or product-extraction calls per minute before you start seeing 429s, while lighter endpoints keep their full 300/min budget.
Rejected requests do not consume budget — if a weighted call is throttled, none of the 10 units are charged and other requests in the same window are unaffected.
The
X-RateLimit-Remaining header reflects the post-weight count, so a successful /web/crawl on a fresh 300/min window drops X-RateLimit-Remaining from 300 to 290.
Read rate-limit headers on every response
Every response from an authenticated request now includes three headers so you can pace requests without waiting for a 429:
Use
X-RateLimit-Remaining to slow down proactively — for example, add a small delay when it drops under 10% of X-RateLimit-Limit — instead of retrying after a 429.
What a 429 looks like
The API returns a JSON envelope:credits_consumed is always 0 on a 429 — throttled requests are never charged.
Every 429 response also includes the standard X-RateLimit-* headers plus a Retry-After header with the precise number of seconds (1–60) until your per-minute window resets:
status === 429. The SDK does not retry automatically. You wire that in.
Pattern 1: Client-side cache for hot domains
The cheapest way to stay under the cap is to skip the call. Brand data changes on the order of months, so a 24-hour client cache is safe for most products:Pattern 2: Backoff on 429 with Retry-After
When you hit rate limits, you get a 429 status code on the response:
Retry-After header tells you exactly how many seconds until your window resets, so use it as the wait time when it’s present. Fall back to exponential backoff (wait 1 second before the first retry and double the delay on each subsequent attempt) if you can’t read the header.
Here’s an example of a retry script that honors Retry-After and falls back to exponential delays:
Pattern 3: Prefetch to shift slow work ahead of bursts
Bursty traffic (like when a marketing email triggers 200 signups in 60 seconds) can get you rate limited. Prefetching doesn’t reduce the number of Brand API calls that count against your limit; every user-facing/brand/retrieve still spends rate-limit budget. What it does is shift the slow crawl work earlier, so each call during the burst completes in under a second instead of stalling for up to a minute and piling up retries on top of an already-saturated window.
Here’s how it works:
- During the burst, your application calls
POST /utility/prefetchright when it first receives the target domain or email, passingtype: "brand"and eitheridentifier.domainoridentifier.email. Prefetch is rate-limit-free, so 200 calls in a minute is fine. - A few seconds later, when the user actually submits and the user-facing client hits the Brand API, the request lands on a warm cache and returns in under a second. That call still counts toward your per-minute limit; it’s just fast.
Pattern 4: Degrade gracefully when the limit holds
If exponential backoff has run out of retries and you are still seeing 429s, the user is better served by a missing-data fallback than an error screen. Some examples:- Onboarding form. Skip the prefilled fields. Let the user enter them by hand and do not block on the API.
- Logo wall. Render the customer’s name in a styled box instead of the logo.
- CRM enrichment. Queue the contact for an offline enrichment job that runs overnight.
Related resources
Prefetch
Warm the cache so burst-time calls return fast.
Best practices
Cache, fallback, and proxy patterns end to end.
Troubleshooting
Other status codes, retry logic, and SDK gotchas.
Pricing
Per-plan credit, rate limit, and overage details.