October 4, 2026 · Ahmed · More by Ahmed
Async Tools, Background Mode, Steering: Not the Same
OpenAI's Responses API has three different non-blocking features. Here's the mechanical difference between async tool calling, background mode, and mid-turn steering.

Alongside GPT-6 Astra's release on September 8, 2026, OpenAI's Responses API picked up two new controls for long-running work — async tool calling and mid-turn steering — joining a third, older one: background mode. All three let a pipeline avoid blocking while something slow happens. That's where the similarity ends. Reach for the wrong one and you'll either block on work that didn't need to block, or build polling logic for something that was never going to run in the background in the first place.
Who actually runs the slow work
Background mode runs response generation itself asynchronously. Set background: true on a Responses API request, and the model's work — the whole response — happens on OpenAI's side without your process holding a connection open. You poll the response by ID while its status is queued or in_progress, or stream it and track a resumable sequence_number cursor so a dropped connection doesn't cost you the work. You can cancel an in-flight background response, and cancelling twice is idempotent — it just returns the final Response object both times.
Async tool calling is a different thing entirely, and the two are easy to conflate because both have "doesn't block" as the headline. Set async: true on a function or custom tool definition, and the model's turn continues after it issues that call instead of pausing to wait for the result. But your application still executes the tool — OpenAI doesn't move execution anywhere or manage a background job for you. When your job finishes, you feed the result back in a later request using the call's original call_id, so the model can match the result to the call that produced it. This only applies to function and custom tools your application runs; it doesn't apply to hosted built-in tools, and it isn't meant to be combined with parallel tool calls in Multi-agent mode.
Mid-turn steering is the third, and it lives over the Responses API's WebSocket mode. A client can send new instructions while a response is still being generated, and the service folds them into a continuation without discarding the work already completed. Continuations chain with previous_response_id the same way HTTP mode does; WebSocket mode adds stream_id on top, which lets one persistent connection run several ordered conversation lanes at once.
Picking the right one for a slow external call
For a commerce-automation pipeline calling slow external APIs — a CRM lookup, an ad-platform call, an AI generation step — the question to ask first is where the slow part actually lives. If the slow part is the model generating a long response, background mode is the fit: you're waiting on OpenAI, so let OpenAI run it async and poll or stream for the result. If the slow part is your own call to a third-party API that the model needs to keep working around, async tool calling is the fit: the model carries on with whatever else it can do while your application's job runs, and you reconcile the result against call_id once it's done. Mid-turn steering solves a third, unrelated problem — letting a client correct or extend an in-progress response — and only applies when you're already on the WebSocket transport.
One pattern worth calling out for async tool calling specifically: a "wait tool." Rather than making the model block on a pending async result by default, you can define an ordinary synchronous tool — no async: true — that the model calls only when it actually needs that pending result. The application binds a task_handle to the original call_id and the running job, and resolves it before returning. That gives the model the choice of whether to wait, instead of your pipeline forcing the wait on every call.
Compatibility is narrower than it looks
Async tool calling is supported by GPT-6 Astra and later models — it isn't a general Responses API feature you can turn on for any model. Combined with the built-in-tools and parallel-tool-call restrictions above, that makes it worth checking model and architecture compatibility before you design a pipeline around it, rather than assuming the three features are interchangeable knobs on the same switch.