PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Build a Multi-Agent System in Python (and When Not To)

Build a coordinator that delegates to specialist AI agents in Python, trace every call, and see what a multi-agent system really costs.

Level
Intermediate
Reading time
24 min
Published
Oct 10, 2026
By
Promptise Team
  • Multi-Agent Systems
  • AI Agents
  • Agent Orchestration
  • MCP
  • Python
  • Promptise Foundry

A multi-agent system splits one job between several AI agents: a coordinator that talks to the user, and specialists that each own one area and its tools. It's the pattern behind most "AI agents talking to each other" demos, and it's easy to build badly. In this guide you'll build a real one in Python: a support coordinator that delegates to a billing specialist and a help center specialist, each with its own MCP server. You'll watch the delegation calls happen, see the coordinator split a question the wrong way and fix it, and measure what the team costs against a single agent with the same tools. Every snippet was run against Promptise Foundry 1.2.1, and every output is what it printed.

[01]

How do you build a multi-agent system in Python?

Build each specialist as an ordinary agent with its own tools. Then build a coordinator and hand it the specialists through cross_agents. With Promptise Foundry, each specialist becomes a tool named ask_agent_<name> that the coordinator's model can call like any other:

Pythonquickstart.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
import asyncio
import sys

from promptise import build_agent
from promptise.config import StdioServerSpec
from promptise.cross_agent import CrossAgent


async def main():
    # A specialist is an ordinary agent with its own tools.
    billing = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": StdioServerSpec(command=sys.executable, args=["billing_server.py"])},
        instructions="You are the billing specialist. Answer from your tools only.",
    )
    # The coordinator gets the specialist as a tool.
    coordinator = await build_agent(
        model="openai:gpt-5-mini",
        servers={},
        instructions="You are a support coordinator. Ask the billing specialist about accounts.",
        cross_agents={
            "billing": CrossAgent(agent=billing, description="Looks up one customer's account."),
        },
    )
    try:
        print("Tools:", coordinator.tool_names)
        result = await coordinator.ainvoke(
            {"messages": [{"role": "user", "content": "What plan is customer acme on?"}]}
        )
        print(result["messages"][-1].content)
    finally:
        await coordinator.shutdown()
        await billing.shutdown()


asyncio.run(main())
Output
Tools: ['ask_agent_billing', 'broadcast_to_agents']
Customer "acme" is on the Team plan.

Details:
- Seats: 12
- Price per seat: $15.00 USD
- Total monthly: $180.00 USD
- Next renewal: 2026-11-01
…

The coordinator has no billing tools. It asked the billing specialist, which looked the account up and answered in its own words. That's the whole mechanism. The rest of this guide builds a two-specialist team properly and puts honest numbers on it. The short version: in our runs the team gave correct answers, and it needed two to three times as many model calls as one agent with the same tools.


[02]

How a coordinator and its specialists work together

Each specialist is a complete agent: its own instructions, its own MCP servers, its own reasoning loop. The coordinator only sees one tool per specialist. When it calls that tool, Promptise runs the specialist from start to finish and returns its final reply as the tool result.

Rendering diagram…

cross_agents always adds two kinds of tools to the coordinator:

Tool

The coordinator's model passes

It gets back

ask_agent_<name>, one per specialist

message, and optionally context and timeout_s

The specialist's final reply, as text

broadcast_to_agents

message, and optionally peers and timeout_s

Each specialist's reply, keyed by name

A few facts from the source that shape how you design a team:

  • A specialist starts fresh every time. It receives the message as a user message, plus context as a system message if the coordinator sends one. It never sees the user's original words or the earlier conversation, only what the coordinator chose to write.

  • Only the final reply comes back. The specialist's tool calls and intermediate steps stay inside it. The coordinator works from prose.

  • Delegation runs in-process. The specialists are Python objects in the same program as the coordinator. To put one on another machine, see the section on delegating across processes below.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry from PyPI. Nothing else is needed for this guide.

  • An API key for a model provider. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.

>_Terminal
pip install promptise
export OPENAI_API_KEY="sk-..."

[04]

Build a support team with two specialists, step by step

The team answers questions about customer accounts. Billing knows each customer's plan, seats and invoices. The help center knows how the product works and what each plan includes and costs. Some questions need one of them, some need both.

Step 01

Give each specialist its own MCP server

Each specialist owns one MCP server. Splitting the tools this way is the main reason to build a team at all: the help center specialist can't read invoices, and the coordinator can't call anything directly.

Pythonbilling_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
from promptise.mcp.server import MCPServer

server = MCPServer("billing")

# A stand-in for your billing system.
ACCOUNTS = {
    "acme": {"plan": "Team", "seats": 12, "price_per_seat": 15.00, "renews": "2026-11-01"},
    "globex": {"plan": "Starter", "seats": 3, "price_per_seat": 8.00, "renews": "2026-10-20"},
}
INVOICES = {
    "acme": [
        {"invoice_id": "INV-2041", "date": "2026-10-01", "amount": 180.00, "status": "paid"},
        {"invoice_id": "INV-1987", "date": "2026-09-01", "amount": 165.00, "status": "paid"},
    ],
    "globex": [
        {"invoice_id": "INV-2044", "date": "2026-10-01", "amount": 24.00, "status": "overdue"},
    ],
}


@server.tool()
async def get_account(customer_id: str) -> dict:
    """Get a customer's plan, number of seats, price per seat in USD and renewal date.

    Args:
        customer_id: The customer's ID, for example "acme".
    """
    account = ACCOUNTS.get(customer_id.strip().lower())
    if account is None:
        return {"error": f"No customer with ID {customer_id}."}
    return {"customer_id": customer_id.strip().lower(), **account}


@server.tool()
async def list_invoices(customer_id: str) -> list[dict]:
    """List a customer's invoices, newest first, with amount in USD and status.

    Args:
        customer_id: The customer's ID, for example "acme".
    """
    return INVOICES.get(customer_id.strip().lower(), [])


if __name__ == "__main__":
    server.run()
Pythondocs_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
from promptise.mcp.server import MCPServer

server = MCPServer("docs")

# A stand-in for your help center or knowledge base.
ARTICLES = [
    {
        "title": "Adding seats",
        "text": "Admins add seats under Settings > Billing > Seats. New seats are billed "
        "immediately, prorated to the end of the current billing period.",
    },
    {
        "title": "Changing your plan",
        "text": "Upgrades take effect immediately and are prorated. Downgrades take effect "
        "at the next renewal date. Plans: Starter (8 USD per seat, up to 5 seats), "
        "Team (15 USD per seat), Business (25 USD per seat, includes SSO).",
    },
    {
        "title": "Overdue invoices",
        "text": "If an invoice is more than 14 days overdue, the workspace becomes read-only "
        "until it is paid. Pay under Settings > Billing > Invoices.",
    },
    {
        "title": "Single sign-on (SSO)",
        "text": "SAML SSO is available on the Business plan. Set it up under Settings > Security.",
    },
]


@server.tool()
async def search_help_center(query: str) -> list[dict]:
    """Search the help center and return matching articles with their full text.

    Args:
        query: A few keywords, for example "seats" or "SSO".
    """
    words = [w for w in query.lower().split() if len(w) > 2]
    return [a for a in ARTICLES if any(w in (a["title"] + " " + a["text"]).lower() for w in words)]


if __name__ == "__main__":
    server.run()

Notice where the Business plan's price lives: in the help center, not in billing. That detail matters in a moment. If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python explains these files line by line.

Step 02

Build the specialists

A specialist is a normal build_agent call. Keep its instructions narrow, and ask for short answers: everything it writes becomes input the coordinator has to read and pay for.

Pythonsupport_team.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
    billing = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": StdioServerSpec(command=sys.executable, args=["billing_server.py"])},
        instructions=(
            "You are the billing specialist. Answer in a few short lines, with facts "
            "from your tools only. Say so if your tools have no answer."
        ),
    )
    docs = await build_agent(
        model="openai:gpt-5-mini",
        servers={"docs": StdioServerSpec(command=sys.executable, args=["docs_server.py"])},
        instructions=(
            "You are the help center specialist. Answer in a few short lines, from help "
            "center articles only, and name the article you used."
        ),
    )

Each specialist can use a different model, a different system prompt, or its own guardrails and approval rules. They're independent agents that happen to be called by another agent.

Step 03

Build the coordinator

The coordinator gets no servers, only its two colleagues. The description on each CrossAgent becomes part of that tool's description, so it's the main thing the model reads when it decides whom to ask:

Pythonsupport_team.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
    coordinator = await build_agent(
        model="openai:gpt-5-mini",
        servers={},
        instructions=(
            "You are a customer support coordinator. You can't see customer data or the "
            "help center yourself, so delegate. Ask the billing specialist about one "
            "customer's account. Ask the docs specialist how the product works, and what "
            "each plan includes and costs. Ask only the specialists the question needs, "
            "with short, specific questions. When it needs both, ask both in the same "
            "step, then combine their answers into one short reply. Don't offer actions "
            "nobody on the team can take."
        ),
        cross_agents={
            "billing": CrossAgent(
                agent=billing,
                description=(
                    "Looks up one customer's account: their plan, seats, what they pay "
                    "today, and their invoices. Doesn't know other plans' prices."
                ),
            ),
            "docs": CrossAgent(
                agent=docs,
                description=(
                    "Answers from the help center: how-to steps, billing policies, and "
                    "what every plan includes and costs per seat."
                ),
            ),
        },
    )

Each description says what the specialist knows and, for billing, what it doesn't. That last sentence wasn't in the first version, and the section after the run shows why it's there.

"Ask both in the same step" is there for speed. When the model asks for two tools in one turn, Promptise runs them in parallel, so the two specialists work at the same time.

Step 04

Trace who asked whom

trace_tools=True prints MCP tool calls, but not calls to ask_agent_* tools, so it can't show you the delegation. LangChain callbacks can: the callbacks you pass to the coordinator's ainvoke also receive every event from the specialists it calls. This small handler uses that to print each delegation, and to count model calls and tokens per agent:

Pythonteam_trace.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
from collections import Counter

from langchain_core.callbacks import BaseCallbackHandler


class TeamTrace(BaseCallbackHandler):
    """Print who asked whom, and count model calls and tokens per agent."""

    def __init__(self, coordinator: str = "coordinator"):
        self.coordinator = coordinator
        self.agent_of_run: dict = {}  # tool run id -> the agent that run belongs to
        self.calls: Counter = Counter()
        self.tokens: Counter = Counter()

    def _agent(self, parent_run_id) -> str:
        return self.agent_of_run.get(parent_run_id, self.coordinator)

    def on_tool_start(self, serialized, input_str, *, run_id, parent_run_id=None, inputs=None, **kwargs):
        name = serialized.get("name", "?")
        caller = self._agent(parent_run_id)
        indent = "" if caller == self.coordinator else "    "
        if name.startswith("ask_agent_"):
            peer = name.removeprefix("ask_agent_")
            self.agent_of_run[run_id] = peer
            print(f"{caller} -> {peer}: {inputs or input_str}")
        else:
            self.agent_of_run[run_id] = caller
            print(f"{indent}{caller} calls {name} {inputs or input_str}")

    def on_tool_end(self, output, *, run_id, parent_run_id=None, **kwargs):
        agent = self.agent_of_run.get(run_id)
        caller = self._agent(parent_run_id)
        if agent != caller:  # a specialist just answered the agent that asked it
            text = str(getattr(output, "content", output))
            print(f"{agent} -> {caller}: {text[:300]}{' …' if len(text) > 300 else ''}")

    def on_llm_end(self, response, *, parent_run_id=None, **kwargs):
        agent = self._agent(parent_run_id)
        usage = response.generations[0][0].message.usage_metadata or {}
        self.calls[agent] += 1
        self.tokens[agent] += usage.get("total_tokens", 0)

    def summary(self) -> str:
        rows = [f"{a}: {self.calls[a]} model calls, {self.tokens[a]:,} tokens" for a in self.calls]
        total = f"total: {sum(self.calls.values())} model calls, {sum(self.tokens.values()):,} tokens"
        return "\n".join(rows + [total])

The trick is the run ID. A specialist's model calls and tool calls arrive with the ask_agent_* call that started them as their parent, so the handler can tell whose work each event is.

Step 05

Ask a question that needs one specialist

The script passes the handler with config={"callbacks": [trace]} on each question. The complete file is at the end of this section.

>_Terminal
python support_team.py
Output
Coordinator tools: ['ask_agent_billing', 'ask_agent_docs', 'broadcast_to_agents']

=== Does customer globex have any unpaid invoices?
coordinator -> billing: {'message': "Please look up customer account for 'globex' and tell me whether they have any unpaid invoices. Include invoice IDs, amounts, due dates, and status for any unpaid invoices. If none unpaid, please confirm that. Provide only the billing info."}
    billing calls get_account {'customer_id': 'globex'}
    billing calls list_invoices {'customer_id': 'globex'}
billing -> coordinator: Unpaid invoices: Yes — 1 overdue invoice.

Invoice ID: INV-2044 | Amount: $24.00 | Invoice date: 2026-10-01 | Due date: not provided by tools | Status: overdue
>>> Yes. Globex has 1 unpaid (overdue) invoice:
…
--- 13.8s
coordinator: 2 model calls, 2,161 tokens
billing: 2 model calls, 1,304 tokens
total: 4 model calls, 3,465 tokens

The coordinator asked only billing, in a sentence of its own making. Billing called both of its tools and answered. Then the coordinator rewrote that answer for the user. Four model calls in total: two to decide and reply, two inside the specialist.

Step 06

Ask a question that needs both

Pricing SSO for acme needs their seat count from billing and the Business plan's price from the help center:

Output
=== Customer acme wants SSO. What would their monthly bill be, and how do they set it up?
coordinator -> billing: {'message': "Lookup customer account 'acme'. Provide: current plan name, number of seats, current monthly amount they pay today, … 'context': None, 'timeout_s': 30}
coordinator -> docs: {'message': 'From the help center: list what each plan includes and the cost per seat for each plan. … 'context': None, 'timeout_s': 30}
    docs calls search_help_center {'query': 'plans includes cost per seat SSO add-on fees setup IdP metadata configure in app testing SCIM user provisioning help center'}
    billing calls get_account {'customer_id': 'acme'}
    billing calls list_invoices {'customer_id': 'acme'}
    docs calls search_help_center {'query': 'SCIM user provisioning SAML IdP metadata Settings > Security setup testing "Single sign-on" help center SAML setup IdP metadata SCIM'}
billing -> coordinator: - Plan: Team
- Seats: 12
- Price per seat: $15.00 USD
- Current monthly amount (today): $180.00 USD (12 × $15.00)
…
docs -> coordinator: - Plans & cost per seat (from "Changing your plan"): Starter — $8/seat (up to 5 seats); Team — $15/seat; Business — $25/seat (Business includes SSO).  
…
>>> Short answer

- Acme is on the Team plan with 12 seats, currently paying $180/month (12 × $15).
- SSO is only included on the Business plan ($25/seat). To get SSO you must upgrade from Team → Business.
- New monthly cost if they upgrade: 12 × $25 = $300/month — an increase of $120/month.

How to enable SSO (high-level)
1. Upgrade the account to the Business plan (Settings → Billing → Plan).  
2. After upgrade, go to Settings → Security → Single sign‑on (SAML).  
3. In your IdP, create a new SAML app and collect IdP metadata (SSO URL, Entity ID, x.509 certificate).  
…
--- 31.3s
coordinator: 2 model calls, 3,983 tokens
docs: 3 model calls, 2,860 tokens
billing: 2 model calls, 1,866 tokens
total: 7 model calls, 8,709 tokens

Both delegations went out in the same step, and the specialists' tool calls interleave because they ran at the same time. The coordinator combined the two answers into the right number, 300 USD a month.

Look closely at two details. First, the coordinator set timeout_s to 30 seconds on its own; the model chooses that argument, not you. More on that below. Second, step 3 of the setup instructions isn't in the help center. The docs specialist said where SSO lives, and the coordinator filled in the IdP steps from general knowledge. Every agent in the chain rewrites what it was told, and every rewrite is a chance to add something that isn't in your data. Instruct the coordinator to pass on only what the specialists said if that matters for your use.

The complete team

Pythonsupport_team.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
import asyncio
import sys
import time

from promptise import build_agent
from promptise.config import StdioServerSpec
from promptise.cross_agent import CrossAgent

from team_trace import TeamTrace

QUESTIONS = [
    "Does customer globex have any unpaid invoices?",
    "Customer acme wants SSO. What would their monthly bill be, and how do they set it up?",
]


async def main():
    # Two specialists, each with its own MCP server and its own instructions.
    billing = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": StdioServerSpec(command=sys.executable, args=["billing_server.py"])},
        instructions=(
            "You are the billing specialist. Answer in a few short lines, with facts "
            "from your tools only. Say so if your tools have no answer."
        ),
    )
    docs = await build_agent(
        model="openai:gpt-5-mini",
        servers={"docs": StdioServerSpec(command=sys.executable, args=["docs_server.py"])},
        instructions=(
            "You are the help center specialist. Answer in a few short lines, from help "
            "center articles only, and name the article you used."
        ),
    )

    # The coordinator has no tools of its own, only its two colleagues.
    coordinator = await build_agent(
        model="openai:gpt-5-mini",
        servers={},
        instructions=(
            "You are a customer support coordinator. You can't see customer data or the "
            "help center yourself, so delegate. Ask the billing specialist about one "
            "customer's account. Ask the docs specialist how the product works, and what "
            "each plan includes and costs. Ask only the specialists the question needs, "
            "with short, specific questions. When it needs both, ask both in the same "
            "step, then combine their answers into one short reply. Don't offer actions "
            "nobody on the team can take."
        ),
        cross_agents={
            "billing": CrossAgent(
                agent=billing,
                description=(
                    "Looks up one customer's account: their plan, seats, what they pay "
                    "today, and their invoices. Doesn't know other plans' prices."
                ),
            ),
            "docs": CrossAgent(
                agent=docs,
                description=(
                    "Answers from the help center: how-to steps, billing policies, and "
                    "what every plan includes and costs per seat."
                ),
            ),
        },
    )

    try:
        print("Coordinator tools:", coordinator.tool_names)
        for question in QUESTIONS:
            print(f"\n=== {question}")
            trace = TeamTrace()
            started = time.monotonic()
            result = await coordinator.ainvoke(
                {"messages": [{"role": "user", "content": question}]},
                config={"callbacks": [trace]},
            )
            print(">>>", result["messages"][-1].content)
            print(f"--- {time.monotonic() - started:.1f}s")
            print(trace.summary())
    finally:
        await coordinator.shutdown()
        await billing.shutdown()
        await docs.shutdown()


asyncio.run(main())

Shut the coordinator down first, then the specialists. Each shutdown() closes that agent's MCP connections and stops its stdio servers.


[05]

When the coordinator splits the question wrong

The first version of this team had vaguer descriptions. Billing was "Looks up a customer's plan, seats, prices and invoices", docs was "Answers how-to and policy questions from the help center", and the instructions said the same. Here's that version on the SSO question:

Output
coordinator -> billing: {'message': "Customer identifier: 'acme'. Please provide:\n- Their current plan name and billing cadence (monthly/annual).\n- Number of active seats/licenses.\n- Current monthly recurring charge (before taxes/discounts).\n- Does enabling SSO change the monthly price or require an add-on? …
coordinator -> docs: {'message': "For customer 'acme' who requested SSO:\n- Which plans include SSO by default, and which require an add-on or upgrade?\n- Supported SSO protocols (SAML, OIDC) and IdPs.\n- Step-by-step setup outline for an admin to enable SSO (short bullets).\n- Required admin permissions and typical timeline/verification steps.\n- Any docs links or KB article IDs to send to …
…
billing -> coordinator: - Plan & cadence: Team — billed monthly (recent invoices on 2026-09-01 and 2026-10-01).
- Active seats: 12.
- Current MRC (before taxes/discounts): $180.00/month (12 × $15.00).
- SSO pricing/add-on: my tools have no information about SSO pricing, add-ons, or proration.
…
>>> Short answer
- Current monthly bill (acme): $180/month (Team plan, 12 seats × $15) — invoices up to date.
- SSO: Not included on Team. SSO is available on the Business plan (SAML SSO). That means acme must upgrade to Business to get SSO.
- I don’t have the Business-plan monthly price or any SSO add-on/proration details from billing yet, so I can’t give the exact new monthly total until I get that quote.
…

The help center had the Business price all along. The coordinator asked billing for the price because the description said billing knew "prices", and asked docs about plans, protocols and setup steps. Billing honestly said it didn't know, and the price fell between the two. The coordinator split the question before it knew where each fact lived, and nothing in the team corrected that.

The fix was to say where each fact lives: billing knows "what they pay today" and "doesn't know other plans' prices"; docs knows "what every plan includes and costs per seat". Describing docs more broadly had a side effect, though. On the next run the coordinator asked docs about globex's overdue invoice too, which cost three extra model calls, so the final version adds "Ask only the specialists the question needs." Routing in a team is prompt engineering, and each change moves it both ways. Re-run your real questions after every change to a description.


[06]

Should you use multiple agents? What it costs

To find out, the same two questions went to the team and to one agent connected to both MCP servers, three times each. The single agent is the same build_agent call with both servers:

Pythoncompare.py
1
2
3
4
5
6
7
8
9
10
11
async def build_single():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": BILLING, "docs": DOCS},
        instructions=(
            "You are a customer support assistant. Use the billing tools for a customer's "
            "plan, seats, prices and invoices, and the help center for how the product "
            "works and what each plan includes and costs. Answer in one short reply."
        ),
    )
    return agent, [agent]
Output
team         | one specialist   | median: 4 calls, 3,409 tokens, 10.8s
…
team         | both specialists | 6 calls |  6,771 tokens |  20.9s | asked ['billing', 'docs'] | correct True
team         | both specialists | 6 calls |  8,888 tokens |  23.8s | asked ['billing', 'docs'] | correct True
team         | both specialists | 7 calls |  8,102 tokens |  27.1s | asked ['billing', 'docs'] | correct True
team         | both specialists | median: 6 calls, 8,102 tokens, 23.8s
…
single agent | one specialist   | median: 2 calls, 1,007 tokens, 3.2s
…
single agent | both specialists | median: 2 calls, 1,822 tokens, 7.6s

Question

Design

Model calls

Tokens

Time

Unpaid invoices (one specialist)

Team

4

3,409

10.8 s

Unpaid invoices (one specialist)

Single agent

2

1,007

3.2 s

SSO price (both specialists)

Team

6

8,102

23.8 s

SSO price (both specialists)

Single agent

2

1,822

7.6 s

Medians of three runs each, with gpt-5-mini. All twelve answers were correct. For this job the team cost three to four and a half times the tokens and took about three times as long, and bought nothing in return. The single agent asked for all the tools it needed in one step, read the raw data itself, and answered. The team paid for a coordinator turn to plan, a full agent run per specialist, prose written by each specialist, and a coordinator turn to read and rewrite that prose.

So start with one agent and more tools. Reach for a team when one of these is true:

  • The tools need different access. Each specialist holds only its own server's credentials, and the coordinator can't call any tool directly. In the run above, the coordinator's whole toolset was ['ask_agent_billing', 'ask_agent_docs', 'broadcast_to_agents'].

  • One agent would have too many tools. A specialist with five tools chooses better than a generalist with eighty.

  • The jobs need different models or rules. A cheap model for lookups, a stronger one for analysis, or approval and guardrails on one specialist only.

  • Different teams own the parts. A billing team can ship and test its specialist without touching the coordinator, and several coordinators can share it.

If none of those apply, a team is extra latency and a second place for answers to go wrong.


[07]

Timeouts, failures and loops

Timeouts are the model's choice unless you set one

timeout_s on ask_agent_* is an argument the coordinator's model fills in. Across the runs for this guide it chose 30 seconds nine times, 10 seconds twice, and often left it out. When it runs out, the tool returns a string instead of raising:

Output
1. ask with timeout_s=1: 'Timed out waiting for peer agent reply.' 1.1s
1b. broadcast with timeout_s=1: {'slow': 'Timed out'} 1.0s

A specialist that legitimately needs 40 seconds will be cut off by a model that guessed 30. For a limit you control, set max_invocation_time on the specialist itself. When it runs out, the coordinator receives the error as its tool result and can tell the user:

Pythonprobe_timeouts.py
1
2
3
4
5
6
    billing = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": StdioServerSpec(command=sys.executable, args=["billing_server.py"])},
        instructions="You are the billing specialist.",
        max_invocation_time=1,  # far too short, on purpose
    )
Output
3. coordinator's tool result: Error: TimeoutError: Agent invocation exceeded 1s timeout
3. coordinator replied: Sorry — I tried to ask our billing specialist but the request failed. I can't access billing data myself.
…

Any exception inside a specialist reaches the coordinator the same way, as a tool result that starts with Error:. Setting max_invocation_time on the coordinator caps the whole team, specialists included. With a specialist that never answered:

Output
4. coordinator raised: TimeoutError Agent invocation exceeded 15s timeout after 15.0s

Nothing limits delegation depth

There's no depth limit and no loop detection when agents run. In practice you rarely build a loop by accident, because CrossAgent takes an agent that already exists, so the graph you build in Python can only point backwards. But if you wire one up, for example with a proxy so an agent can reach itself, nothing stops it. In a test where an agent always delegated back into itself, a guard of our own that refused after four levels bounded the depth, but not the retries: the agent at each level tried again when refused.

Output
>>> Depth limit reached. Answer it yourself.
50 model calls, 135.4s

That was one question, "What is 2 + 2?". Keep the delegation graph a tree, and set max_invocation_time on the top agent as a backstop. If you define agents in .superagent files, the loader does catch circular file references: each cross_agents entry there is a CrossAgentConfig with a file and a description, and two files that refer to each other fail to load with Circular cross-agent reference detected.

Turn off the broadcast tool if you don't want it

broadcast_to_agents sends the same message to every specialist. That suits "ask everyone for a second opinion", not a team where each specialist has a different job, and it doesn't pass context. To leave it out, build the ask tools yourself with make_cross_agent_tools and pass them as extra_tools:

Pythonprobe_no_broadcast.py
        extra_tools=make_cross_agent_tools(
            {"billing": CrossAgent(agent=billing, description="Looks up one customer's account.")},
            include_broadcast=False,
        ),
Output
Coordinator tools: ['ask_agent_billing']

[08]

Delegate to an agent in another process

cross_agents only takes agents in the same Python process. Promptise's older network server, which served agents to each other over HTTP, has been removed. To run a specialist on another machine, or let another team own it, put it behind an MCP server: one tool that takes a question and returns the specialist's answer.

Pythonbilling_agent_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
import sys

from promptise import build_agent
from promptise.config import StdioServerSpec
from promptise.mcp.server import MCPServer

server = MCPServer("billing-team")
specialist = {}


@server.on_startup
async def start_specialist():
    specialist["agent"] = await build_agent(
        model="openai:gpt-5-mini",
        servers={"billing": StdioServerSpec(command=sys.executable, args=["billing_server.py"])},
        instructions=(
            "You are the billing specialist. Answer in a few short lines, with facts "
            "from your tools only. Say so if your tools have no answer."
        ),
    )


@server.on_shutdown
async def stop_specialist():
    await specialist["agent"].shutdown()


@server.tool()
async def ask_billing_specialist(question: str) -> str:
    """Ask the billing specialist about one customer's account: their plan, seats,
    what they pay today, and their invoices. It doesn't know other plans' prices.

    Args:
        question: One short, specific question that names the customer.
    """
    result = await specialist["agent"].ainvoke({"messages": [{"role": "user", "content": question}]})
    return result["messages"][-1].content


if __name__ == "__main__":
    server.run(transport="http", host="127.0.0.1", port=8260)

The specialist is built once when the server starts and shut down when it stops. The coordinator connects to it like any MCP server, and can mix it with in-process specialists:

Pythonremote_coordinator.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
    coordinator = await build_agent(
        model="openai:gpt-5-mini",
        # The billing specialist runs in another process, behind an MCP server.
        servers={"billing_team": HTTPServerSpec(url="http://127.0.0.1:8260/mcp")},
        instructions=(
            "You are a customer support coordinator. Delegate: ask the billing specialist "
            "about one customer's account, and the docs specialist how the product works "
            "and what each plan includes and costs. When a question needs both, ask both "
            "in the same step, then combine their answers into one short reply."
        ),
        cross_agents={
            "docs": CrossAgent(
                agent=docs,
                description=(
                    "Answers from the help center: how-to steps, billing policies, and "
                    "what every plan includes and costs per seat."
                ),
            ),
        },
    )
Output
Coordinator tools: ['ask_billing_specialist', 'ask_agent_docs', 'broadcast_to_agents']
coordinator calls ask_billing_specialist {'question': "Please provide Acme's account details: current plan, number of seats, what they pay today (monthly), and recent invoices."}
coordinator -> docs: {'message': 'Which plans include SSO? For each plan, state the per-seat monthly price and what SSO-related features or setup are included (e.g., SAML, SCIM, provisioning, number of SSO-enabled seats). Also list per-seat monthly prices for all plans.', 'context': None, 'timeout_s': 10}
    docs calls search_help_center {'query': 'SSO plans include SSO per-seat monthly price SAML SCIM provisioning number of SSO-enabled seats per-seat prices for all plans'}
…
>>> Short answer: to give Acme SSO they'd need the Business plan at $25/seat/month. Acme currently on Team: 12 seats × $15 = $180/month today. If upgraded to Business: 12 × $25 = $300/month — an increase of $120/month.  
…
coordinator: 2 model calls, 2,512 tokens
docs: 2 model calls, 1,690 tokens
total: 4 model calls, 4,202 tokens

The remote specialist's model calls don't appear in the count: they happen in the server's process, so the coordinator's callbacks never see them. Count them on the server if you need the full cost. Before you expose a specialist like this on a network, require an API key or a token, as Authentication & Security shows. A specialist behind a tool acts on whatever question it's sent, within the limits of its own tools.


[09]

Honest limits

These are true of Promptise Foundry 1.2.1:

  • Delegation is in-process only. Use an MCP server, as above, to reach a specialist elsewhere.

  • Some older examples don't run. The Multi-Agent Coordination page passes URL strings to cross_agents and calls agent.ask_peer() and agent.broadcast(). None of that exists in 1.2.1; a URL string fails with AttributeError: 'str' object has no attribute 'description'. Use CrossAgent as in this guide.

  • The model picks `timeout_s`. Set max_invocation_time on each specialist and on the coordinator for limits you control.

  • No depth limit or loop detection at runtime. Keep the graph a tree.

  • Specialists see only the coordinator's message. No conversation history, no user identity in the text. Put what a specialist needs into the message or context.

  • `trace_tools` doesn't show delegation. Use a callback handler like TeamTrace, or Promptise's observability.

  • The `promptise agent` command can't run a team whose specialists use stdio servers. It builds the specialists from .superagent files in one event loop and runs the coordinator in another, so every specialist tool call failed with Not connected in our test. Build teams in Python until that's fixed.


[10]

Frequently asked questions

When should you use a multi-agent system instead of a single agent?

When the parts need different access, different models, or different owners, or when one agent would face too many tools to choose well. Otherwise use one agent with more tools. In this guide's measurements, the single agent answered the same questions correctly with a third to a quarter of the tokens, in about a third of the time.

What is multi-agent orchestration?

Deciding which agent does which part of a job and combining the results. This guide uses the most common pattern, a coordinator (also called a supervisor) that delegates to specialists through tool calls. Promptise also supports agents that react to each other's events through the Agent Runtime, which suits long-running pipelines better than a request and reply.

How do AI agents talk to each other?

In this pattern, through tool calls. The coordinator's model writes a message, the specialist runs and replies in text, and that text comes back as the tool's result. They don't share memory or conversation history, so everything a specialist needs has to be in the message.

Do I need CrewAI or LangGraph to build a multi-agent system in Python?

No. A specialist is an agent and delegation is a tool call, so any agent framework with tool calling can do it. With Promptise Foundry it's the cross_agents argument of build_agent, and each specialist keeps everything an agent has: MCP servers, approval, guardrails and memory.

Can the specialists use different models?

Yes. Each specialist is its own build_agent call with its own model string, so a lookup specialist can run on a small, cheap model while the coordinator uses a stronger one.


[11]

Where to go next

  • Cross-Agent Delegation: CrossAgent, the generated tools and their arguments.

  • Cross-Agent API reference: make_cross_agent_tools and its options.

  • SuperAgent files: define agents and their cross_agents in YAML.

  • Building Agents: every build_agent option, including max_invocation_time.

  • Agent Runtime: long-running agents with triggers and events.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: give a specialist your existing API as tools.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths