The direct answer
Tool use (also called function calling) is the mechanism that lets an AI model reach outside its own text output. Normally a model just writes words. With tool use, you hand it a list of software functions it may call — look up a database record, fetch the weather, issue a refund, run a search — and the model decides, mid-reply, whether it needs one. If it does, it sends back a structured request naming the tool and its inputs.
The model never runs the software itself. Your code executes the function, hands the result back to the model, and the model folds it into its answer. That loop — model asks, code acts, model answers — is what turns a chatbot into something that can actually do things in software.
How it works
The loop has five steps. First, you send your request to the model with the tools included: each tool is defined by a name, a plain-language description, and a schema for its inputs (for example, a get_weather tool that takes a location). Second, the model replies — not with text, but with a tool call: the tool's name plus the arguments it wants to use. Third, your code runs the real function with those arguments. Fourth, you send the tool's output back to the model. Fifth, the model writes its final answer — or makes another tool call, and the loop continues.
Providers split tools two ways. Client tools are ones you define and execute yourself, like a refund function wired into your order system. Server tools run on the provider's infrastructure, like web search or sandboxed code execution that Anthropic or OpenAI runs for you. One practical detail: every tool definition rides along as input tokens, so a huge menu of tools adds cost to every request — a reason teams keep the menu small or load rarely used tools only when the model reaches for them.
A simple example
As an illustration, imagine a travel assistant. You ask: "What's the weather in Denver this weekend?" The model sees a tool named get_weather that accepts a location, and it responds with a tool call: get_weather(location="Denver, CO"). Your code calls a real weather API, gets back "sunny, 72°F", and feeds that to the model. The model then answers: "Sunny and 72°F this weekend in Denver."
Notice what each party did. The model did the deciding — which tool, which arguments. Your code did the doing — the actual API call. The model then did the explaining. Deciding, doing, explaining: that division of labor is the entire concept.
Why it matters
Without tool use, a model is frozen at its training cutoff, blind to your private data, and unable to act on anything. Tool use is the difference between a chatbot that talks about your order and one that can find and refund it. It connects the model's reasoning to the real world: live prices, your calendar, your codebase, your customers.
It's also the foundation of AI agents — systems that loop through plan, call tools, read results, and adjust, over many steps, to complete a goal. Every agent you've heard about is tool use wearing a trench coat: the same request-call-result loop, run repeatedly with a plan on top.
The common misunderstanding
The model is not running code or browsing the web by itself.
It only emits a structured request — a name and some arguments. All actual execution happens in your code or on the provider's servers. That means a tool call is exactly as safe as the code behind it: if you hand a model a delete_database tool with no confirmation step, the resulting disaster is an engineering failure, not a model misbehaving. Tool use doesn't give the model hands; it gives you a way to lend it yours — so design the tools as carefully as you'd design any API.
What changed recently
Both major ecosystems keep expanding what a tool can be. OpenAI's documentation now describes schema-defined functions alongside free-text custom tools and a set of built-in tools for web search, code execution, and connecting to MCP servers. Anthropic's docs draw a clean line between client tools (you execute) and server tools (Anthropic executes, results returned with citations).
The other trend is deferred loading: when an application has hundreds of possible tools, newer APIs let the model search for and load only the relevant ones mid-conversation, instead of stuffing the entire menu into every request. Tool use is getting both more powerful and more economical at the same time.
Try it on PlainLogic
Agent Mission on PlainLogic puts you in the orchestrator's seat: you hand an AI a goal and a set of tools, then watch it plan, call tools, and handle the results — the same loop this article describes.
Sources
- PRIMARY SOURCEOpenAI: Function calling guide
- PRIMARY SOURCEAnthropic: Tool use with Claude