PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Too Many MCP Tools? LLM Tool Selection, Measured

What 122 tools do to an LLM agent's accuracy, tokens and latency, measured, and how semantic tool selection fixes it in Python, with honest limits.

Level
Intermediate
Reading time
18 min
Published
Oct 10, 2026
By
Promptise Team
  • Tool Selection
  • MCP
  • AI Agents
  • Token Optimization
  • Python
  • Promptise Foundry

Connect a couple of MCP servers to an agent and you can easily end up with 100 or more tools. Every one of them goes to the model on every call, as a name, a description and a JSON schema. That costs tokens and time, and past a provider's limit the request is rejected outright. This guide measures what too many tools really does to an LLM agent. You'll build a realistic 122-tool SaaS admin MCP server, give an agent 20 admin tasks with known right answers, and compare all tools against each of Promptise Foundry's tool optimization levels: right tool chosen, input tokens and latency. Every snippet ran against Promptise Foundry 1.2.1, and the results table comes from 200 recorded agent runs.

[01]

What happens when an LLM agent has too many tools?

Three things. Every model call carries every tool definition, so cost and latency grow with the tool count: 122 tools came to about 11,000 input tokens per call before the model read a word of the task. Past a limit the call fails: OpenAI rejects requests with more than 128 tools. And look-alike tools (suspend_user and delete_user, create_domain and create_dns_record) give the model more ways to pick the wrong one. The fix is to send each request only the tools that fit it. With Promptise Foundry, that's one argument:

Pythonfollowup.py
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"platform": StdioServerSpec(command=sys.executable, args=["platform_server.py"])},
        optimize_tools="semantic",  # each request gets only the tools that fit it
    )

Before each request, the agent compares the user's message with every tool description using a small local embedding model, and offers the model the 8 best matches. Here's one of the recorded runs, "dep_9f2c is building the wrong branch. Stop it.":

Output
[semantic] task 18, expected cancel_deployment
  offered: ['cancel_deployment', 'delete_project', 'archive_project', 'resize_database', 'delete_deployment', 'request_quota_increase', 'rollback_deployment', 'resolve_incident']
  called:  ['cancel_deployment']
  answer:  Stopped building deployment dep_9f2c.

Across the experiment, the median first call dropped from about 11,000 input tokens to about 700. It isn't free, though. In this experiment, semantic selection with default settings left the right tool out of the 8 for 3 of the 20 tasks. The rest of this guide shows the measurements, where it fails, and how to tune it.


[02]

How tool optimization works

optimize_tools works in two layers. The static layer runs once, when the agent is built: it strips parameter descriptions from the schemas and shortens long tool descriptions. The semantic layer runs on every request:

Rendering diagram…

The three presets, from Promptise's source (promptise/tool_optimization.py):

Preset

Parameter descriptions

Tool descriptions cut at

Nested schemas

Tools per request

"minimal" (also True)

Removed

200 characters

Kept

All

"standard"

Removed, nested ones too

150 characters

Flattened beyond depth 3

All

"semantic"

Removed, nested ones too

100 characters

Flattened beyond depth 2

8

Two details matter later. The text that gets embedded is name: description after shortening, so the semantic preset matches requests against the first 100 characters of each description. And selection happens once per ainvoke(), from the last message only. The model keeps those 8 tools for the whole run.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry and sentence-transformers, which semantic selection uses for embeddings. It isn't installed with Promptise; without it, optimize_tools="semantic" fails in build_agent with ModuleNotFoundError: No module named 'sentence_transformers'. It pulls in PyTorch, so expect a large download. This guide used sentence-transformers 6.1.0.

  • An OpenAI API key. The default embedding model, all-MiniLM-L6-v2, runs locally. Promptise downloads it from Hugging Face on first use and caches it in ~/.cache/huggingface/.

>_Terminal
pip install promptise sentence-transformers
export OPENAI_API_KEY="sk-..."

[04]

Measure tool selection on 122 tools, step by step

Step 01

Build an MCP server with too many tools

You need a server that looks like the real thing: many resources, the usual list, get, create, update and delete for each, plus the special actions every admin API grows. Writing 122 functions by hand would be silly, so this helper generates them from tables. Each tool just echoes its arguments back, which is all you need to see which one the model picked:

Pythonacme_tools.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
def make_tool(server, name: str, description: str, params: dict[str, str]):
    """Register one tool. Parameters whose description starts with "Optional" may be left out."""

    async def handler(**kwargs):
        return {"ok": True, "tool": name, "arguments": kwargs}

    parameters, annotations = [], {}
    for p, d in params.items():
        optional = d.startswith("Optional")
        annotation = Annotated[str | None if optional else str, Field(description=d)]
        parameters.append(inspect.Parameter(p, inspect.Parameter.KEYWORD_ONLY, annotation=annotation,
                                            default=None if optional else inspect.Parameter.empty))
        annotations[p] = annotation
    handler.__signature__ = inspect.Signature(parameters)
    handler.__annotations__ = annotations
    handler.__name__ = name
    server.tool(name=name, description=description)(handler)


def add_tools(server, resources: dict, actions: dict):
    """Five standard tools per resource, plus one tool per action."""
    for res, (plural, example, what, fields) in resources.items():
        noun = res.replace("_", " ")
        id_param = {f"{res}_id": f'The {noun} ID, for example "{example}".'}
        make_tool(server, f"list_{res}s",
                  f"List {plural} in the organization, newest first. A {noun} is {what}. "
                  f"Returns IDs you can pass to get_{res}.",
                  {"filter": "Optional text to filter by.", "limit": "Optional maximum number of results."})

The server file is mostly data: 14 resources and 52 actions, each with descriptions written the way a careful API team would, including when to use one tool rather than its look-alike:

Pythonplatform_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
server = MCPServer("acme-platform")

# resource: (plural, ID example, what it is, fields you can set on create and update)
RESOURCES = {
    "user": ("users", "u_42", "a person with a login in your organization",
             {"email": "Work email address.", "role": "One of owner, admin, member, billing or viewer."}),
}

# Actions beyond list, get, create, update and delete. name: (description, {parameter: description})
ACTIONS = {
    "suspend_user": ("Suspend a user so they can't sign in. Their data and memberships are kept and can be restored. "
                     "Use this, not delete_user, when someone leaves the company.",
                     {"user_id": "The user to suspend.", "reason": "Why, recorded in the audit log."}),
    "list_audit_events": ("Search the audit log: who did what, and when, across the organization. "
                          "Use it to find out who changed or deleted something.",
                          {"action": "Action to filter by, for example database.delete.", "since_days": "How far back."}),
}

add_tools(server, RESOURCES, ACTIONS)

The complete files are in the evidence for this guide. A second server, customer_server.py, holds 26 billing and support tools; you'll need it in a moment.

Step 02

Measure what the model receives

Before spending anything on model calls, look at the payload. This script builds an agent at each level and estimates the tokens of the tool definitions, serialized the way LangChain sends them to OpenAI. It's an offline estimate; the API's own count, in Step 5, came out lower.

Pythondefinitions.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
"""How big are the tool definitions the model receives, at each optimization level?"""
import asyncio
import json
import sys

import tiktoken
from langchain_core.utils.function_calling import convert_to_openai_tool

from promptise import build_agent
from promptise.config import StdioServerSpec

SERVERS = {"platform": StdioServerSpec(command=sys.executable, args=["platform_server.py"])}
enc = tiktoken.get_encoding("o200k_base")


def definition_tokens(tools) -> int:
    payload = [convert_to_openai_tool(t) for t in tools]
    return len(enc.encode(json.dumps(payload)))


async def main():
    for level in [None, "minimal", "standard", "semantic"]:
        agent = await build_agent(model="openai:gpt-5-mini", servers=SERVERS, optimize_tools=level)
        try:
            tools = agent.tools
            suspend = next(t for t in tools if t.name == "suspend_user")
            print(f"--- optimize_tools={level!r}: {len(tools)} tools, {definition_tokens(tools)} tokens")
            print(json.dumps(convert_to_openai_tool(suspend)["function"]))
        finally:
            await agent.shutdown()


asyncio.run(main())
Output
--- optimize_tools=None: 122 tools, 12977 tokens
{"name": "suspend_user", "description": "Suspend a user so they can't sign in. Their data and memberships are kept and can be restored. Use this, not delete_user, when someone leaves the company.", "parameters": {"properties": {"user_id": {"description": "The user to suspend.", "type": "string"}, "reason": {"description": "Why, recorded in the audit log.", "type": "string"}}, "required": ["user_id", "reason"], "type": "object"}}
--- optimize_tools='minimal': 122 tools, 10585 tokens
{"name": "suspend_user", "description": "Suspend a user so they can't sign in. Their data and memberships are kept and can be restored. Use this, not delete_user, when someone leaves the company.", "parameters": {"properties": {"user_id": {"type": "string"}, "reason": {"type": "string"}}, "required": ["user_id", "reason"], "type": "object"}}
--- optimize_tools='standard': 122 tools, 10542 tokens
{"name": "suspend_user", "description": "Suspend a user so they can't sign in. Their data and memberships are kept and can be restored. Use this, not delete_user, when someone leaves the...", "parameters": {"properties": {"user_id": {"type": "string"}, "reason": {"type": "string"}}, "required": ["user_id", "reason"], "type": "object"}}
--- optimize_tools='semantic': 123 tools, 10122 tokens
{"name": "suspend_user", "description": "Suspend a user so they can't sign in. Their data and memberships are kept and can be restored....", "parameters": {"properties": {"user_id": {"type": "string"}, "reason": {"type": "string"}}, "required": ["user_id", "reason"], "type": "object"}}

Watch what happens to suspend_user. minimal drops the parameter descriptions. standard cuts the description mid-sentence. semantic cuts it before the one sentence that says when to use it. The semantic line counts all 123 tools (122 plus a fallback tool), but the model only ever sees 8 of them per request, as Step 4 shows.

Note

The docs estimate about 40% savings for minimal on tool definitions. On this server it was 18%, because these schemas are flat and their parameter descriptions are short. Servers with long, nested schemas save more.

Step 03

Write tasks with one right answer

Twenty requests, phrased the way people actually ask, each with the one tool that does the job. Several are deliberately close to a dangerous look-alike:

Pythontasks.py
1
2
3
4
5
6
7
8
9
TASKS = [
    ("Jane (u_42) left the company yesterday. Make sure she can't get in any more.", "suspend_user"),
    ("The last deploy of prj_web_shop broke checkout. Put production back the way it was.", "rollback_deployment"),
    ("Our CI key key_8812 leaked in a public repo. Give us a fresh secret for it with the same permissions.", "rotate_api_key"),
    ("Point www.acme.io at acme-shop.pages.dev with a CNAME.", "create_dns_record"),
    ("Somebody removed db_staging last week. Find out who.", "list_audit_events"),
    ("dep_9f2c is building the wrong branch. Stop it.", "cancel_deployment"),
    ("How much was key_8812 used this week?", "get_api_key_usage"),
]
Step 04

Record tools, tokens and time for every model call

The experiment builds one agent per setting and runs every task against each, interleaved so they share the same network conditions. A LangChain callback records what each model call was given and what it cost. invocation_params holds the tools actually sent, so for semantic selection you see the 8 that were chosen, not the full list:

Pythonexperiment.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
SETTINGS = {
    "all tools": None,
    "minimal": "minimal",
    "standard": "standard",
    "semantic": "semantic",
    "semantic, full descriptions": ToolOptimizationConfig(
        level=OptimizationLevel.SEMANTIC, max_description_length=400
    ),
}


class ModelCalls(AsyncCallbackHandler):
    """Records the tools offered, token usage and duration of every model call."""

    def __init__(self):
        self.calls = []

    async def on_chat_model_start(self, serialized, messages, **kwargs):
        tools = (kwargs.get("invocation_params") or {}).get("tools") or []
        self.calls.append({"tools": [t["function"]["name"] for t in tools], "start": time.perf_counter()})

    async def on_llm_end(self, response, **kwargs):
        call = self.calls[-1]
        usage = response.generations[0][0].message.usage_metadata or {}
        call["seconds"] = time.perf_counter() - call.pop("start")
        call["input_tokens"] = usage.get("input_tokens", 0)
        call["cached_tokens"] = (usage.get("input_token_details") or {}).get("cache_read", 0)

Pass it per request with agent.ainvoke(..., config={"callbacks": [recorder]}). A task counts as solved when the expected tool was called at any point in the run.

Step 05

Read the results

Each of the 20 tasks ran twice per setting with openai:gpt-5-mini: 40 runs per setting, 200 in total. A third repeat was planned but stopped when the API account ran out of credits, so it's left out. Medians are per task; tokens are what the OpenAI API reported.

Setting

Right tool offered

Right tool called

Tools per call

Input tokens, first call

Cached by OpenAI

Seconds per task

Seconds, first call

All tools

40/40

38/40 (95%)

122

11,067

89%

3.6

2.0

minimal

40/40

37/40 (92%)

122

9,471

91%

4.0

2.5

standard

40/40

37/40 (92%)

122

9,428

91%

4.6

2.5

semantic

34/40

32/40 (80%)

8

691

0%

2.9

1.6

semantic, full descriptions

38/40

36/40 (90%)

8

692

0%

2.7

1.6

What the numbers say:

  • gpt-5-mini handled 122 tools well. With every tool in view it got 38 of 40 right. Too many tools didn't break accuracy here; it made every call carry 11,000 tokens.

  • The static presets saved 14% of input tokens and changed little else. The one or two runs between all tools, minimal and standard are within noise at 40 runs each.

  • Semantic selection cut first-call input by 94%, from 11,067 to 691 tokens, and both semantic settings were the fastest. The preset also lost the most tasks: six of its eight misses happened because the right tool wasn't among the 8 offered.

  • Caching narrows the cost gap. OpenAI served 89% of the all-tools input from its prompt cache, because the 122 definitions are the same at the start of every request. Cached tokens are billed at a discount, so the real cost difference is smaller than the raw token counts suggest. The semantic prompts got no cache hits at all: they're short, and OpenAI only caches prompts above a minimum length.

Two tasks were hard for most settings, and they aren't about tool count. "Point www.acme.io at acme-shop.pages.dev with a CNAME" sent the all-tools agent to create_domain once and to a clarifying question once: the request really is ambiguous between the two tools. "Bring back last night's backup" failed in 8 of 10 runs because this test server's list_database_backups returns no backups to choose from, so the agent kept looking instead of restoring. That's the test, not the agent.


[05]

Where semantic tool selection hurts

Where semantic selection lost a task the other settings solved, it lost it before the model saw anything: the right tool wasn't offered. Here's the first task under both semantic settings, from the same recording:

Output
[semantic] task 1, expected suspend_user
  offered: ['list_incidents', 'remove_team_member', 'delete_team', 'list_users', 'create_incident', 'delete_incident', 'delete_webhook', 'add_team_member']
  called:  ['list_users']
  answer:  I can’t disable or delete a user account with the tools available here, so I can’t fully lock Jane (u_42) out of the org by myself. …

[semantic, full descriptions] task 1, expected suspend_user
  offered: ['suspend_user', 'remove_team_member', 'list_incidents', 'list_users', 'add_team_member', 'delete_team', 'list_webhook_deliveries', 'acknowledge_incident']
  called:  ['suspend_user']
  answer:  Done — Jane (u_42) has been suspended and can no longer sign in.

"Jane left the company" doesn't look much like "Suspend a user so they can't sign in". The sentence that does match, "when someone leaves the company", sits after character 100, and the semantic preset cuts it off before embedding. Keep the full description and suspend_user comes first.

Two things went right in the failures, and they matter. The model never grabbed a destructive substitute: offered delete_team instead of suspend_user, it said it couldn't do the job. And it said so plainly, so a person would know to step in.

The audit question failed under both semantic settings. "Somebody removed db_staging last week. Find out who" is about deleting databases, word for word, so the index offered delete_database first and list_audit_events not at all. Embedding models match topics, not intent. The model didn't call delete_database, but it was one bad decision away.

Test retrieval before you pay for model calls

Every miss above can be found without calling the model. The agent's index is ToolIndex in promptise.tool_optimization. It isn't on the docs page, but it's what build_agent uses. Build it over the agent's own tools and check each task's expected tool against the top 8:

Pythonretrieval_check.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
"""Before you pay for a single model call: would semantic selection hand the model the right tool?

Builds the same index the agent builds (same tools, same truncated descriptions,
same embedding model) and checks each task's expected tool against the top 8.
"""
import asyncio
import sys

from promptise import OptimizationLevel, ToolOptimizationConfig, build_agent
from promptise.config import StdioServerSpec
from promptise.tool_optimization import ToolIndex

from tasks import TASKS

SERVERS = {"platform": StdioServerSpec(command=sys.executable, args=["platform_server.py"])}

SETUPS = {
    "semantic preset": (ToolOptimizationConfig(level=OptimizationLevel.SEMANTIC), "all-MiniLM-L6-v2"),
    "full descriptions": (ToolOptimizationConfig(level=OptimizationLevel.SEMANTIC, max_description_length=400),
                          "all-MiniLM-L6-v2"),
    "full descriptions, bge-small": (ToolOptimizationConfig(level=OptimizationLevel.SEMANTIC, max_description_length=400),
                                     "BAAI/bge-small-en-v1.5"),
}


async def main():
    for label, (config, embedding_model) in SETUPS.items():
        agent = await build_agent(model="openai:gpt-5-mini", servers=SERVERS, optimize_tools=config)
        try:
            tools = [t for t in agent.tools if t.name != "request_more_tools"]
            index = ToolIndex(tools, model_name_or_path=embedding_model)
            misses = []
            for text, expected in TASKS:
                names = [t.name for t in index.select(text, top_k=8)]
                if expected not in names:
                    misses.append(f"{expected} (got {', '.join(names[:3])}, ...)")
            print(f"{label}: right tool in the top 8 for {len(TASKS) - len(misses)}/{len(TASKS)} tasks")
            for miss in misses:
                print("  missed", miss)
        finally:
            await agent.shutdown()


asyncio.run(main())
Output
semantic preset: right tool in the top 8 for 17/20 tasks
  missed suspend_user (got list_incidents, remove_team_member, delete_team, ...)
  missed create_dns_record (got request_quota_increase, verify_domain, create_domain, ...)
  missed list_audit_events (got delete_database, delete_team, resize_database, ...)
full descriptions: right tool in the top 8 for 19/20 tasks
  missed list_audit_events (got delete_database, resize_database, promote_deployment, ...)
full descriptions, bge-small: right tool in the top 8 for 19/20 tasks
  missed create_dns_record (got request_quota_increase, search_logs, get_organization, ...)

This offline check predicted the live runs exactly: 17 of 20 for the preset, 19 of 20 with full descriptions, the same misses, at no model cost. Swapping the embedding model for BAAI/bge-small-en-v1.5 fixed the audit question and broke the DNS one, so a different model isn't automatically a better one. Run your own real requests through this check before you switch semantic selection on.

How to tune it in 1.2.1

  • Keep full descriptions. ToolOptimizationConfig(level=OptimizationLevel.SEMANTIC, max_description_length=400) lifted accuracy from 80% to 90% here, with the same input tokens per call.

  • Put the "when to use it" first. The index reads name: description. A description that opens with the situations people describe ("someone left", "find out who changed something") matches requests better than one that opens with what the endpoint does.

  • Try another embedding model with embedding_model=, a Hugging Face name or a local path, and keep it only if your retrieval check improves.

  • Don't count on `semantic_top_k`, `preserve_tools` or `request_more_tools` in 1.2.1. They don't reach the per-request selection; see the honest limits below.


[06]

Past 128 tools, all-tools stops working

The first version of this experiment put all 148 tools, platform and customer, in one server. With all tools sent, OpenAI refused every call:

Output
{"error": "GraphExecutionError: Graph 'react' failed at node 'reason': OpenAIInvalidRequestError: Error code: 400 - {'error': {'message': \"Invalid 'tools': array too long. Expected an array with maximum length 128, but got an array with length 148 instead.\", 'type': 'invalid_request_error', 'param': 'tools', 'code': 'array_above_max_length'}}", "setting": "all tools", "rep": 0, "task": 1, "expected": "suspend_user"}

minimal and standard failed the same way, since they still send every tool. That's why the measured server above has 122. If your agent connects several MCP servers, count the total: two ordinary servers can cross 128 together. Semantic selection always sends 8 tools, so it never hits this limit. OpenAI's own function calling guide suggests far fewer, under 20 at the start of a turn, as a soft rule of thumb.


[07]

Follow-up messages get the wrong tools

Selection reads only the latest message. That works for one-shot requests and breaks in a conversation. Here are the tools offered for the same job, asked directly and then confirmed as a follow-up:

Output
One message: ['retry_webhook_delivery', 'rollback_deployment', 'promote_deployment', 'delete_project', 'cancel_deployment', 'list_webhook_deliveries', 'delete_deployment', 'delete_webhook']
Follow-up:   ['acknowledge_incident', 'create_database_backup', 'delete_incident', 'request_quota_increase', 'list_databases', 'add_team_member', 'remove_team_member', 'delete_team']

After "Yes, go ahead.", rollback_deployment is gone, and the model can't do what it just offered to do. If your agent chats, either keep semantic selection for single-shot jobs, or put the full request into the latest message before you call it.


[08]

The other fix: fewer, better tools

Selection works around a big tool list. The other strategy is to not have one. Many MCP servers are a one-to-one copy of a REST API, so they inherit dozens of look-alike CRUD endpoints. MCPcast, Promptise's OpenAPI-to-MCP generator, takes the opposite approach: it drops what agents shouldn't touch and can merge related endpoints into one well-named tool. On the Swagger Petstore it turned 19 operations into 12 tools. OpenAPI to MCP: Turn Any REST API into an MCP Server walks through it, with an evaluation of how well a real agent uses the result.

The two combine well. Curate the servers you own; use semantic selection for the ones you don't, or when the total is still over a hundred.


[09]

Honest limits

These are true of Promptise Foundry 1.2.1 and were checked against its source and a run:

  • `semantic_top_k` is ignored. The agent calls the index without it, so the model always gets 8 tools. With semantic_top_k=3 and semantic_top_k=20, our probe printed model saw 8 tools both times.

  • `preserve_tools` doesn't keep a tool in the selection. The docs say preserved tools are always selected; in practice the preserved get_organization was missing from the 8. It also doesn't keep parameter descriptions: those are stripped while the schema is converted, before the preserve list is checked. It does keep the tool description from being shortened.

  • The `request_more_tools` fallback never reaches the model. It's added to the agent's tool list, but not to the per-request selection, so across all 80 semantic runs the model never saw it. When the index misses, the model can't recover on its own.

  • Selection happens once per request, from the last message. Multi-step jobs need every tool in the first 8, and follow-ups get tools chosen from text like "Yes, go ahead."

  • Shortened descriptions end in odd ways, for example can be restored...., because the cut adds ... after the sentence's own full stop. Harmless, but visible to the model.

  • This is one model, one server and 20 tasks. Two runs per task per setting is enough to see a 10-point gap, not a 2-point one. Measure with your own tools and requests before you decide.


[10]

Frequently asked questions

How many tools can an LLM handle?

It depends on the model and the tools. gpt-5-mini picked the right tool in 95% of runs with 122 tools in view, so the ceiling is usually set by cost, latency and provider limits before accuracy. OpenAI rejects more than 128 tools in one request, and recommends far fewer as a rule of thumb.

How many tools should an LLM agent see?

As few as the task needs. In this experiment, 8 tools per request cut input tokens by 94%, but the wrong 8 cost up to 15 points of accuracy. Aim for the smallest set that still contains the right tool, and check that with your real requests, as in the retrieval check above.

What is semantic tool selection?

Choosing which tools to offer the model by embedding similarity. The tool descriptions are embedded once, each request is embedded when it arrives, and only the closest matches go to the model. It's also called tool retrieval or dynamic tool selection. In Promptise Foundry it's build_agent(optimize_tools="semantic"), and it runs locally with sentence-transformers.

Does tool selection work with any MCP server?

Yes. Optimization happens in the agent, after it lists the servers' tools, so the servers don't need to change. It works with any tool-calling model, because the model just receives a shorter tool list.

Is it better to merge tools or to filter them?

Merge when you own the server: fewer, clearer tools help every client, not just your agent. Filter when you don't own the server, or when you connect so many that even good tools add up to more than a hundred.


[11]

Where to go next

  • Tool Optimization: every ToolOptimizationConfig option and preset.

  • Building Agents: the rest of build_agent.

  • Observability: track token usage across your agents.

  • MCPcast reference: generate a curated MCP server from an OpenAPI spec.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: fewer, better tools in practice.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths