← Back to concepts
8 min read

Tool Calling

Tool calling is the mechanism that lets a model ask an application to run a function on its behalf, then use the result to keep working. It is one of the main building blocks for agents. Without tools, the model cannot fetch fresh data or take actions by itself. The application can still place information in the prompt, but the model has no direct reach beyond that context.

For example, if a user asks “What is the weather in Paris right now?”, the model cannot know this from training alone. Tool calling lets it request a get_weather(city) function, receive the answer, and use it to reply.

How a tool call works

A tool call usually follows this sequence:

  1. The application describes the available tools to the model — their names, purposes, and inputs.
  2. The model decides a tool is needed and the API returns a structured tool-call object, usually with a call ID, function name, and JSON arguments.
  3. The application executes the actual function or API call.
  4. The result is sent back to the model as part of the conversation.
  5. The model reads the result and continues, either answering the user or calling another tool.

The model never runs code itself. It only produces a request. The application is responsible for executing it safely and returning the result.

A tool call from model request to final answer The user asks a question. Inside the application host, the model receives the request with tool definitions and returns a structured tool request with a name and JSON arguments. The application validates the request, executes the tool or API outside the model, and sends the result back as a new message. The model reads that result and can answer the user or choose another call. The model never executes the tool directly. User question Application host Model chooses next step Structured request get_weather { "city": "Paris" } Application validates and executes Tool / API runs outside Result new message
The model only proposes a structured call; the application validates it, executes the tool, and returns the result for the model to read.

Describing tools to the model

Each tool needs a clear definition so the model knows when and how to use it. A typical tool definition includes:

  • Name: a short, unambiguous identifier, like search_orders.
  • Description: what the tool does and when to use it.
  • Parameters: the expected inputs, usually as a structured schema (name, type, required or optional).

For example:

Tool: search_orders
Description: Find a customer's orders by email or order ID.
Parameters:
  - query (string, required): email address or order ID

Vague descriptions lead to misuse. If two tools look similar, the model may call the wrong one or pass bad arguments.

Good tool design makes the next step obvious:

  • Use clear names and descriptions that say when to use the tool and when not to use it.
  • Use JSON Schema with required fields, types, ranges, and enums where choices are limited.
  • Keep tools small and focused. A refund_order tool is easier to secure than a broad run_admin_action tool.
  • Return concise results with the facts the model needs next, not a full raw API dump.
  • Return helpful errors the model can act on, such as order_not_found or missing_email.
  • Make write actions idempotent when possible, so retries do not duplicate work.
  • Set timeouts and retry only safe failures.

Structured output for calls

Tool calls are usually returned as structured data rather than free text. Modern APIs expose them as tool-call objects. The application should not scrape prose to guess what the model meant.

{
  "id": "call_42",
  "name": "search_orders",
  "arguments": { "query": "user@example.com" }
}

Structured calls reduce ambiguity. The application can read the call ID, function name, and arguments directly, validate them, and send the matching result back under the same ID. This is related to prompt engineering and structured output, but the important point is that tool calls are machine-readable requests, not ordinary assistant text.

Tool definitions turn text into a checkable request A tool definition lists a name, description, and parameter schema. The model produces JSON with a call id, the tool name search_orders, and arguments containing a query. The application validates that the tool exists, required fields are present, and the argument types are correct. Valid requests execute; invalid requests return a clear error message to the model. Tool definition name: search_orders description: find orders parameters: query: string required: true Model output { "id": "call_42", "name": "search_orders", "arguments": { "query": "user@example.com" } } Application validation tool name exists required query present query is a string valid: execute invalid: clear error
A structured tool call is useful because the application can validate it before anything runs.

Handling the result

Once a tool runs, the host appends its output to the conversation as tool-result context. The model reads that result on the next turn. It has not learned the information into its weights; it is simply using new context for this interaction. The result should be clear, relevant, and shaped for the next model step.

If a tool call fails — for example, the order was not found — the error should also be returned to the model in a clear form. A silent failure can cause the model to guess or hallucinate an answer instead of reporting the problem.

Parallel and multi-step calls

Some tasks need more than one tool call. If the calls are independent, the model may request them in one assistant turn and the host can run them in parallel. For example, it can fetch weather and hotel availability for the same city at the same time.

Dependent calls must happen in sequence. If the second call needs an ID returned by the first call, the host should wait, append the first result, and let the model choose the next call. Each result is matched to its call ID so the model can tell which output belongs to which request.

Parallel tool calls return matched results One assistant turn contains two independent tool calls, call A for weather and call B for hotels. The host validates both calls, runs the tools in parallel, and returns two tool results with the same call IDs. The model reads both results before continuing. A separate note says dependent calls should run in sequence. Assistant turn 2 independent tool calls call_A get_weather call_B search_hotels Weather tool runs now Hotel tool runs now Results matched by ID If call_B needs call_A's output, run them in sequence instead.
Parallel calls save time only when the calls do not depend on each other's results.

Why tool calling needs guardrails

Tool calling gives a model real-world reach, so it also introduces risk. A model could call a tool with the wrong arguments, call a destructive tool by mistake, or be tricked into misusing a tool through malicious input in the conversation.

Common safeguards include:

  • Narrow, single-purpose tools instead of broad, powerful ones.
  • Authorizing every call as the end user, not as “the model.”
  • Validating arguments server-side before execution.
  • Requiring human confirmation for sensitive actions, like deleting data or sending money.
  • Treating tool output as untrusted data that may contain prompt injection, especially when it came from web pages, files, tickets, or emails.
  • Logging every call and its result for audit and debugging.
  • Limiting call rates, budgets, and how many calls can happen in a row.

The host or harness enforces these safeguards. The model can request a call, but it should not be the authority that decides whether the user is allowed to run it.

Guardrails around a tool call A user message and model request enter the application boundary. The boundary applies five checks: narrow tool choice, argument validation, confirmation for sensitive actions, call limits, and logging. Safe read-only calls execute immediately. Sensitive actions pause for human confirmation. Invalid or over-limit calls are blocked and returned as errors instead of reaching the external system. Request user + model call Application guardrails narrow tool choice argument validation confirmation if sensitive call count and cost limits log call and result Execute safe call read-only search or lookup Pause for approval delete, send, spend, publish blocked: invalid, unsafe, or over limit
Guardrails keep tool calling useful by making execution conditional, observable, and reversible where possible.

Tool calling vs MCP

Tool calling is the general mechanism: a model requests a function, the application runs it, and the result comes back. MCP is a standard protocol for exposing tools (along with resources and prompts) in a consistent format across many applications. You can have tool calling without MCP, but MCP builds on the same core idea.

The key idea

Tool calling lets a model request actions instead of only generating text. The model proposes a call, the application executes it safely, and the result flows back into the conversation. Clear tool definitions, structured requests, and safety checks are what make tool calling dependable rather than risky.