The OpenAI Agents SDK in Practice: Agents, Handoffs and Guardrails

The OpenAI Agents SDK runs the agent loop for you: an agent calls tools, hands off to another agent and is checked by guardrails, with sessions and tracing built in. Here are the primitives, a working Python example with one tool and one handoff, and when to use the plain Responses API instead.

7 min read

The OpenAI Agents SDK is a Python library (with a TypeScript twin) that runs the agent loop for you. You describe agents, each with instructions and tools; the SDK calls the model, runs the tools it asks for, switches agents when one hands off to another, checks input and output with guardrails, and records every step as a trace. You install it with pip install openai-agents. If you would rather own the loop yourself, the plain Responses API is still there, and OpenAI’s own guidance on choosing between them is short and clear. Both are below, with a small example you can run.

The primitives

The Agents SDK documentation (opens in a new tab) describes a deliberately small set of building blocks. Everything else in the library is built from these:

  • Agent: a model with a name, instructions and tools, plus optional handoffs, guardrails and a structured output type.
  • Tools: Python functions wrapped with @function_tool, whose schema is built from the type hints and docstring; hosted tools such as web search, file search and code interpreter that run on OpenAI’s side; and other agents exposed as tools with Agent.as_tool().
  • Handoffs: one agent passing the conversation to another. The model sees each handoff as a tool named transfer_to_<agent_name>.
  • Guardrails: checks on the user’s input, the final output, or individual tool calls. A guardrail that trips raises an exception and stops the run.
  • Sessions: stored conversation history, so the next run picks up where the last one ended without you passing the transcript back in.
  • Tracing: a record of every model call, tool call, handoff and guardrail in a run, on by default.

The Runner ties them together. Runner.run() is asynchronous, Runner.run_sync() blocks, and Runner.run_streamed() yields events as they happen.

A minimal Python agent: one tool, one handoff

A support agent that can look up an order and passes refund requests to a second agent. It also carries one input guardrail, so you can see where each primitive sits. It needs Python 3.10 or later and an OPENAI_API_KEY in the environment.

Terminal
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install openai-agents
export OPENAI_API_KEY=...        # Windows PowerShell: $env:OPENAI_API_KEY = "..."
support.py
import asyncio

from agents import (
    Agent,
    GuardrailFunctionOutput,
    Runner,
    SQLiteSession,
    function_tool,
    input_guardrail,
)

ORDERS = {"A1001": "shipped", "A1002": "awaiting payment"}


@function_tool
def order_status(order_id: str) -> str:
    """Look up the status of an order.

    Args:
        order_id: The order reference, for example A1001.
    """
    return ORDERS.get(order_id, "no such order")


refunds_agent = Agent(
    name="Refunds",
    handoff_description="Handles refund requests.",
    instructions="Explain the refund policy. Never promise a refund; say a person will confirm it.",
)


@input_guardrail
async def no_card_numbers(ctx, agent, user_input) -> GuardrailFunctionOutput:
    text = user_input if isinstance(user_input, str) else str(user_input)
    digits = sum(ch.isdigit() for ch in text)
    return GuardrailFunctionOutput(output_info={"digits": digits}, tripwire_triggered=digits >= 13)


support_agent = Agent(
    name="Support",
    instructions="Answer order questions with the order_status tool. Hand refund requests to Refunds.",
    tools=[order_status],
    handoffs=[refunds_agent],
    input_guardrails=[no_card_numbers],
)


async def main() -> None:
    session = SQLiteSession("customer-42")
    result = await Runner.run(support_agent, "Where is order A1002?", session=session)
    print(result.final_output)
    print("Answered by:", result.last_agent.name)


if __name__ == "__main__":
    asyncio.run(main())

This was checked against version 0.22.3 of the openai-agents package: every name imports from agents, the file passes mypy, the tool’s generated schema has one required string, order_id, described from the docstring, and the handoff appears to the model as transfer_to_refunds. Ask “I want my money back for A1001” and the run should end with last_agent set to Refunds.

What happens when it runs

  1. The runner calls the model for the current agent with the current input. Because a session is attached, the stored history is put in front of that input first.
  2. The input guardrail runs. By default it runs in parallel with the agent, which is faster but means the model may already have used some tokens before a tripwire stops it; blocking mode waits for the guardrail first.
  3. If the model asks for a tool, the runner calls order_status, appends the result and loops.
  4. If the model calls transfer_to_refunds, the runner switches the current agent to Refunds and loops again. By default the new agent sees the whole conversation; an input_filter on handoff() can trim it.
  5. When the model returns text with no tool calls, that is the final output. The new items are saved to the session and the trace is closed.

A run that keeps calling tools stops at max_turns with a MaxTurnsExceeded exception. Keep that limit in place; it is the cheapest protection you have against a loop.

Handoffs or agents as tools

The handoffs guide (opens in a new tab) draws the line. A handoff transfers control: the receiving agent takes over the conversation and answers the user. Agent.as_tool() keeps control with the calling agent and uses the specialist like any other function, returning its result. Use handoffs when a different agent should own the rest of the exchange, such as refunds or billing; use agents as tools when one coordinator needs several specialists’ answers and writes the reply itself. handoff() also takes on_handoff for a callback and input_type when you want the model to send a reason or priority with the transfer.

Guardrails, and what they cannot do

According to the guardrails page (opens in a new tab), input guardrails run only on the first agent in a run and output guardrails only on the last, so a guardrail on Refunds’ input would never fire in the example above. Tool guardrails, declared with @tool_input_guardrail and @tool_output_guardrail, run on every call to the function tool they wrap, which makes them the right place to catch a secret in an argument or a result. A tripwire raises InputGuardrailTripwireTriggered or OutputGuardrailTripwireTriggered, and your code decides what the user sees.

Guardrails check content; they do not grant or remove permissions. For a tool that changes something, set needs_approval=True on @function_tool. The run then pauses with the pending call in result.interruptions, and you approve or reject it on the run state before resuming.

Sessions and tracing

SQLiteSession is the lightweight option for development, in memory or backed by a file. The sessions documentation (opens in a new tab) lists others for production, including Redis, SQLAlchemy and MongoDB, an encrypted wrapper, and OpenAIConversationsSession, which keeps history on OpenAI’s side. A session holds history in your code, so it cannot be combined in the same run with previous_response_id or conversation_id; pick one mechanism per run.

Tracing is on by default and sends traces to the Traces dashboard on the OpenAI platform. The tracing page (opens in a new tab) gives three ways to turn it off: OPENAI_AGENTS_DISABLE_TRACING=1, set_tracing_disabled(True), or tracing_disabled in a run’s configuration. It also notes that tracing is unavailable to organisations using OpenAI’s APIs under Zero Data Retention, and you can add your own trace processors to send spans elsewhere.

The Agents SDK vs the Responses API

The SDK calls the Responses API underneath for OpenAI models, so this is not a choice between two engines. It is a choice about who owns the loop. OpenAI’s documentation puts it this way:

  • Use the Agents SDK when you want the runtime to manage turns, tool execution, guardrails, handoffs or sessions, when the agent works across several coordinated steps, or when you need resumable execution.
  • Use the Responses API directly when you want to own the loop, tool dispatch and state yourself, or when the job is short-lived and mainly about returning the model’s response.
  • Many applications use both: the SDK for the managed workflow, direct Responses API calls for the simple paths.

OpenAI’s guide to building agents (opens in a new tab) now lists two more options beside these: ChatKit for an embedded chat interface, and an Agents API that runs an agent on OpenAI’s infrastructure with the Codex harness, for long-running work you do not want to host. The SDK is not tied to OpenAI models either; a model provider or an adapter such as LiteLLM lets each agent use a different vendor’s model.

When the SDK is the right tool

  • Several agents with different instructions or tools, and the model should decide who handles what: handoffs.
  • A tool call that must wait for a person: needs_approval.
  • A conversation that spans many requests: sessions.
  • You need to see why a run went wrong: traces, without writing any logging first.
  • Not worth it: one prompt and one answer, or a fixed pipeline where your code already decides every step. A plain API call is easier to test.

MCP servers plug in the same way as tools; which classes to use, and when OpenAI rather than your code makes the call, is covered in MCP with OpenAI models. If you are weighing the SDK against Anthropic’s equivalent, the Claude Agent SDK takes the same questions from the other side.

Give the agent somewhere to report

A trace tells you what happened inside a run. It is not where the rest of the team looks. A useful pattern is to give the agent a board: it reads the task it was given, and when it stops it comments with what it did and moves the task along, or files a bug for what it could not do. fenbs is an MCP server, so an Agents SDK program can reach it with a token you issue by hand under Settings, with a name, the scopes it needs (read, write, comment) and an optional expiry. Everything it changes is recorded in the board’s history under the assistant’s name, and revoking the token stops it at once.

Related

Build a first agent end to end: how to build an AI agent. How agents and tasks divide the work in other frameworks: agents vs tasks. The visual route and its shutdown: OpenAI Agent Builder. Connecting an OpenAI client to a board: the ChatGPT integration.

Questions people ask.

Is the OpenAI Agents SDK free to use?

The library is open source and installs with pip install openai-agents. You pay for the model calls it makes, at the normal API rates of whichever provider you use.

What is the difference between a handoff and an agent used as a tool?

A handoff passes control: the receiving agent takes over the conversation and replies. An agent used as a tool, through Agent.as_tool, is called like a function and returns its answer to the calling agent, which stays in charge.

Can the OpenAI Agents SDK use models from other providers?

Yes. You can set a model provider for a run, give each agent its own model, or use the LiteLLM or Any-LLM adapters. Providers without the Responses API need the Chat Completions model class instead.

How do I stop the Agents SDK sending traces to OpenAI?

Set the environment variable OPENAI_AGENTS_DISABLE_TRACING to 1, call set_tracing_disabled(True), or set tracing_disabled in the run configuration for a single run.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.