← All articles

October 1, 2026 · Ahmed · More by Ahmed

OpenAI Now Splits Rate Limits From Server Overload

OpenAI's API now separates slow_down 429s from server_is_overloaded 503s — and automation pipelines that treat them the same waste retries or get throttled.

OpenAI Now Splits Rate Limits From Server Overload — ASSOSIATIX Journal

Two failure modes that used to be indistinguishable

Alongside GPT-6 Astra's September 2026 release, OpenAI changed how its API reports failure under load. A request that's sent too quickly now returns a 429 with a slow_down code. Temporary overload on OpenAI's own infrastructure returns a different error entirely: a 503 with a server_is_overloaded code. OpenAI's own API changelog and release-notes page confirm the same wording for both, though the two pages disagree on the exact day the change shipped — the API changelog groups it in a September 10 cluster, the release notes page dates the GPT-6 Astra entry September 3. Only the month is solid.

Before this split, a pipeline calling the OpenAI API at volume — batch jobs, bulk content generation, scheduled agents — had to treat "too many requests" and "the model is overloaded" as the same problem, because the API reported them the same way. They aren't the same problem, and treating them identically has a real cost in both directions.

Why conflating them wastes retries either way

A slow_down 429 means your own traffic ramped up too fast — the fix is to reduce request rate or concurrency on your end. A server_is_overloaded 503 means OpenAI's infrastructure is temporarily strained, a condition your pipeline can't fix by throttling; the only lever is to wait and retry. A pipeline that throttles on every error, including genuine server overload, slows itself down for a problem it didn't cause. One that retries both at full speed keeps hammering an endpoint when the real fix was to slow down — and risks turning a self-inflicted rate-limit problem into a longer backoff cycle.

Building retry logic around the two codes

Both error responses can include a Retry-After header. OpenAI's guidance: when it's present, wait at least as long as it specifies before retrying; fall back to exponential backoff only when the header is absent. No specific backoff formula or duration is given beyond that. The practical split for a pipeline's error handling: a slow_down 429 should trigger client-side throttling — reduce concurrency going forward, not just delay the next call. A server_is_overloaded 503 should trigger a wait-and-retry without touching the pipeline's steady-state request rate.

This shipped alongside new Responses API controls built for long-running work — async tool calling, mid-turn steering, and changing reasoning effort mid-conversation — all aimed at exactly the kind of extended, scheduled, or agentic workloads where retry behavior under load compounds over hours, not seconds. Handling the two error codes distinctly is a small change with an outsized effect on throughput for anything running at scale.

Sources

Automation & AI

ASSOSIATIX works on this every day. See our Shopify Operations Automation service

$./start-project.sh

READY TO BUILD
WHAT'S NEXT?

Tell us where you are and where you want to go. We'll help you choose the right path—storefront, system, product, or a focused growth sprint.