AI, usefully · 4 min read ·

How AI uses tools—and who actually runs them

When an AI assistant checks an order, searches a database or creates a calendar event, what is actually happening?

The language model isn’t directly reaching into those systems. In a typical tool-calling setup, it requests an operation. Software around the model decides whether to execute it, runs it and returns the result.

Understanding that separation makes agent orchestration much easier to follow. It also explains where permissions, validation and reliability belong.

What is a tool?

A tool is a function the application makes available to the model.

It might search documents, retrieve an order, calculate a metric or send a message. Each tool has a name, a description and an expected set of inputs.

For example:

{
  "name": "lookup_order",
  "description": "Retrieve an order’s current delivery status.",
  "inputs": {
    "order_id": "string"
  }
}

The description helps the model decide when to use it. The input specification tells it what information to supply. The actual implementation—connecting to the order database—lives in the application.

Tool design matters. Ambiguous names and descriptions make correct selection harder. Anthropic’s guidance recommends clearly defined tools, useful responses and evaluation against representative tasks.

Follow one request through the system

Suppose the user asks:

Where is order 104?

The model sees the question and the available tool definitions. It can respond with a structured request:

{
  "tool": "lookup_order",
  "arguments": {
    "order_id": "104"
  }
}

This is a tool call: a request to run a function with particular arguments. It is not evidence that anything has happened yet.

The application then checks whether the request is valid. Does the tool exist? Is the identifier correctly formatted? Is the signed-in user allowed to access that order?

Only after those checks does it run the lookup.

The tool might return:

{
  "order_id": "104",
  "status": "in_transit",
  "estimated_delivery": "2026-10-04",
  "updated_at": "2026-10-02T15:00:00Z"
}

The application passes that result back to the model. Now the model can explain that the order is in transit, with an estimated delivery date of October 4.

The estimate should remain an estimate in the final answer. Tool access gives the model evidence; it doesn’t guarantee that the model will describe that evidence accurately.

Where orchestration enters

Orchestration is the coordination around these steps: maintaining the task’s state, routing tool requests, handling results and deciding when to continue or stop.

The tool-calling loop
User request
Model selects next step
Tool requested?

Yes

Application checks access and inputs
Execute tool or return an error

Result returns to the model ↺

No

Return answer or ask for clarification

After receiving a tool result or error, the model selects its next step again.

A tool result may answer the question immediately. It may also reveal missing information that requires another lookup.

The application needs limits on this loop. Otherwise, the model could repeatedly retry a failing tool, make unnecessary calls or spend too long pursuing an answer.

Anthropic distinguishes predefined workflows from agents that dynamically choose their steps and tools. Both can use tools; tool calling alone does not make a system autonomous.

Permission belongs in the application

Consider two tools:

ToolWhat it can do
lookup_orderRead delivery information
cancel_orderChange an order’s state

Their consequences are different.

A prompt saying “be careful with cancellations” is insufficient protection. The application should enforce the rules: which orders the user can access, whether cancellation is allowed and whether confirmation is required.

The user’s identity should come from the authenticated session. The application should not trust a model-supplied claim that the requester owns an order.

For an action requiring approval, confirmation should cover the specific action and target—not a vague instruction to “go ahead.”

Check execution and interpretation separately

There are several ways this workflow can fail.

The model can choose the wrong tool. It can supply the wrong identifier. The tool can return outdated information. The model can then misread a perfectly valid result.

A useful test checks each stage:

Test failures deliberately. Try an unknown order, a timeout and an order belonging to another account. The system should handle each case without inventing a delivery status.

For tools that change data, retries need additional care. If an operation succeeded but its response was lost, repeating it could perform the action twice. That requires protection in the application, not simply another instruction to the model.

Try this before adding more agents

Define one read-only tool for a task you understand.

Write down its purpose, required inputs, return fields and expected failures. Then ask an AI assistant:

Given this tool definition and these five user requests, identify when you would call the tool, which arguments you would supply and when you would ask for clarification. Do not invent missing identifiers.

Inspect the proposed calls yourself.

This exercise simulates tool selection; it does not connect the assistant to a real system. But it exposes unclear definitions before you build the execution layer.

Once that single loop is understandable and testable, you have a foundation for supervisors, specialized agents and more complex workflows.