← All articles

September 28, 2026 · Ahmed · More by Ahmed

How OpenAI's Tool Search Defers Tool Definitions at Runtime

OpenAI's tool_search lets agents load only the function definitions they need at runtime, instead of every tool up front.

As agent-style automations pile up tools — CRM lookups, billing actions, order systems, MCP servers — every tool definition loaded up front counts against the context window and gets billed as input tokens on every turn. OpenAI's tool_search, available in the Responses and Agents APIs on gpt-5.4 and later models, lets the model discover and load only the tool definitions it actually needs at runtime instead of front-loading the whole catalog.

How Deferred Loading Works

Activating it takes two steps: add tool_search as a tool in the request's tools array, and mark the functions you want deferred with defer_loading: true (for MCP servers, set the flag on the server's tool definition instead). For namespaces, defer_loading applies to the functions inside the namespace, not the namespace object itself. At the start of a request the model still sees the name and description of whatever is searchable — for an individual deferred function that means it sees the name and description up front, so in practice tool search mostly defers the parameter schema. When the model loads a deferred tool, the definition is injected at the end of the context window specifically so the model's cache is preserved across requests.

Hosted vs. Client-Executed Search

There are two ways to run the discovery step. In hosted tool search, OpenAI searches across the deferred tools declared in the request and returns the loaded subset in the same response — the simplest path when the candidate tools are already known when the request is built. In client-executed tool search, the model emits a tool_search_call, the application performs the lookup itself, and returns a matching tool_search_output. OpenAI recommends starting with hosted search and reaching for client-executed search only when which tools are available depends on tenant, project, or permission state that your application controls, not the request itself.

Where the Limits Sit

Two constraints matter for anyone scoping this: only gpt-5.4 and later models support tool_search in the Responses API, and OpenAI recommends deferring at the namespace or MCP-server level rather than individual functions, since models are primarily trained to search those surfaces and token savings are more material there. The guidance is to keep each namespace under roughly 10 functions and to aim for fewer than 20 functions available at the start of a turn overall. In the Agents API, function definitions load eagerly by default — deferring requires explicitly adding tool_search to agent.tools plus defer_loading: true on each function, and mixing eager and deferred functions in one session is supported but "generally not recommended." MCP tools are the exception: the Agents API discovers them automatically when the model and provider support tool search, with no per-server flag needed.

For teams building out a large tool catalog — the kind that accumulates once a CRM, a billing system, and a handful of MCP servers are all wired into one agent — tool_search is the mechanism for keeping that catalog from eating the context window on every turn, provided the model in use is gpt-5.4 or newer.

Sources

Automation & AI

ASSOSIATIX works on this every day. See our Shopify Operations Automation service

$./start-project.sh

READY TO BUILD
WHAT'S NEXT?

Tell us where you are and where you want to go. We'll help you choose the right path—storefront, system, product, or a focused growth sprint.