PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Best LLM for AI Agents: Test and Switch Models in Python

Find the best LLM for your AI agent with a tool-calling eval, then switch between OpenAI, Anthropic, Azure, Bedrock, Gemini and Ollama in one line.

Level
Intermediate
Reading time
22 min
Published
Oct 10, 2026
By
Promptise Team
  • LLM
  • AI Agents
  • Tool Calling
  • Evaluation
  • Python
  • Promptise Foundry

Every few weeks a new model claims to be the best LLM for AI agents, and every time you wonder whether to switch. The honest answer is that a leaderboard can't tell you, because it didn't run your tools. In this guide you'll make the model a single setting, switch an agent between OpenAI, Anthropic, Azure OpenAI, Bedrock, Gemini, Ollama and any OpenAI-compatible endpoint without touching the agent code, and build a 10-task tool-calling eval that reports accuracy, tokens and latency. Every snippet ran against Promptise Foundry 1.2.1. Only an OpenAI key was available, so the eval compares five OpenAI models with real numbers, and the other providers are configured and checked but not run.

[01]

What is the best LLM for AI agents?

The best LLM for your agent is the cheapest, fastest model that passes your own tool-calling tasks, and the only way to find it is to run those tasks against each candidate. In this guide's eval of 10 shop tasks, run three times each, gpt-4.1-mini, gpt-5-mini and gpt-5 all scored 30 of 30, and gpt-4.1-mini answered in about half the time. The two nano models dropped 2 and 5 answers.

To make that comparison cheap, keep the model out of your agent code. With Promptise Foundry, the model is one argument to build_agent, so read it from configuration:

Pythonagent.py
1
2
3
4
5
6
7
8
# The only line that knows which model runs the agent.
MODEL = os.environ.get("AGENT_MODEL", "openai:gpt-5-mini")


async def main():
    agent = await build_agent(
        model=MODEL,
        servers={"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"])},
>_Terminal
AGENT_MODEL=openai:gpt-4.1-mini python agent.py
AGENT_MODEL=anthropic:claude-sonnet-4-5 python agent.py

The tools, the instructions and the code around them stay exactly the same. Only the string changes, and the new provider needs its own key.


[02]

How switching models works

A model string has two parts: the provider before the colon and that provider's model name after it. Promptise looks up the provider, reads its credentials from the environment or a .env file, and builds the right client. Your tools live on MCP servers, so they don't care which model calls them.

Rendering diagram…

OpenAI, Azure OpenAI and Anthropic use their own LangChain clients. Every other provider is reached through its OpenAI-compatible endpoint, which is why one pip install promptise covers all of them. These are the providers this guide covers:

Provider

Model string

Environment variables

Route

In this guide

OpenAI

openai:gpt-5-mini

OPENAI_API_KEY

native

Run, with real results

Anthropic

anthropic:claude-sonnet-4-5

ANTHROPIC_API_KEY

native

Configured, not run

Azure OpenAI

azure:<deployment name>

AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_KEY, OPENAI_API_VERSION

native

Configured, not run

Amazon Bedrock

bedrock:<model id>

AWS_DEFAULT_REGION, AWS_BEARER_TOKEN_BEDROCK

OpenAI-compatible

Configured, not run

Google Gemini

gemini:gemini-2.5-pro

GOOGLE_API_KEY or GEMINI_API_KEY

OpenAI-compatible

Configured, not run

Ollama

ollama:llama3.1

none, OLLAMA_HOST optional

OpenAI-compatible

Wiring tested against a stub

Any OpenAI-compatible server

openai:<model> plus an endpoint

OPENAI_BASE_URL, OPENAI_API_KEY

OpenAI-compatible

Wiring tested against a stub

Promptise knows thirteen more, including Groq, Mistral, DeepSeek, OpenRouter and Vertex AI. promptise models list prints all of them with their aliases and which variables you've set.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry from PyPI. Nothing else: every provider in the table works with the core install.

  • An API key for at least one provider. This guide uses OpenAI.

>_Terminal
pip install promptise
export OPENAI_API_KEY="sk-..."

[04]

Compare LLMs on your own tool-calling tasks, step by step

You'll give an agent seven shop tools, write ten questions with known answers, and run them against five models.

Step 01

Put your tools in an MCP server

The eval should exercise your real tools. Here they are a small shop, with fixed data so every question has one right answer. In your project, point the eval at the MCP servers your agent already uses.

Pythonshop_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
from promptise.mcp.server import MCPServer

server = MCPServer("shop")

# A small, fixed dataset, so every eval task has one right answer.
PRODUCTS = {
    "MUG-01": {"name": "Ceramic mug", "price_eur": 12.50, "stock": 42},
    "CUP-02": {"name": "Travel cup", "price_eur": 9.90, "stock": 7},
    "LAMP-03": {"name": "Desk lamp", "price_eur": 39.00, "stock": 15},
    "LAMP-04": {"name": "Floor lamp", "price_eur": 89.00, "stock": 4},
    "DESK-05": {"name": "Standing desk", "price_eur": 499.00, "stock": 3},
}
CUSTOMERS = {
    "C-7": {"name": "Dana Ito", "country": "PT", "tier": "gold"},
    "C-8": {"name": "Lee Park", "country": "DE", "tier": "standard"},
    "C-9": {"name": "Sam Ruiz", "country": "US", "tier": "standard"},
}
ORDERS = {
    "A-1001": {"customer_id": "C-8", "items": [{"sku": "MUG-01", "qty": 2}], "status": "shipped"},
    "A-1002": {"customer_id": "C-7", "items": [{"sku": "LAMP-03", "qty": 1}], "status": "processing"},
    "A-1003": {"customer_id": "C-7", "items": [{"sku": "DESK-05", "qty": 1}], "status": "delivered"},
    "A-1004": {"customer_id": "C-7", "items": [{"sku": "CUP-02", "qty": 3}], "status": "cancelled"},
    "A-1005": {"customer_id": "C-9", "items": [{"sku": "LAMP-04", "qty": 1}, {"sku": "MUG-01", "qty": 1}], "status": "shipped"},
}
EUR_RATES = {"USD": 1.087, "GBP": 0.842, "CHF": 0.937}
SHIPPING = {"DE": (4.90, 1.20), "PT": (6.90, 2.10), "US": (19.00, 4.50)}  # base fee, per kg


@server.tool()
async def get_order(order_id: str) -> dict:
    """Get an order: customer ID, items (SKU and quantity), status and total in EUR."""
    order = ORDERS.get(order_id.strip().upper())
    if order is None:
        return {"error": f"No order {order_id}."}
    total = sum(PRODUCTS[i["sku"]]["price_eur"] * i["qty"] for i in order["items"])
    return {"order_id": order_id.strip().upper(), **order, "total_eur": round(total, 2)}


@server.tool()
async def list_orders(customer_id: str) -> list[str]:
    """List the IDs of every order a customer has placed, including cancelled ones."""
    return [oid for oid, o in ORDERS.items() if o["customer_id"] == customer_id.strip().upper()]


@server.tool()
async def get_customer(customer_id: str) -> dict:
    """Get a customer: name, two-letter country code and loyalty tier."""
    customer = CUSTOMERS.get(customer_id.strip().upper())
    if customer is None:
        return {"error": f"No customer {customer_id}."}
    return {"customer_id": customer_id.strip().upper(), **customer}


@server.tool()
async def get_product(sku: str) -> dict:
    """Get a product by SKU: name, unit price in EUR and units in stock."""
    product = PRODUCTS.get(sku.strip().upper())
    if product is None:
        return {"error": f"No product {sku}."}
    return {"sku": sku.strip().upper(), **product}


@server.tool()
async def search_products(query: str) -> list[dict]:
    """Find products whose name contains the query. Returns SKU and name only."""
    q = query.lower().rstrip("s")
    return [{"sku": s, "name": p["name"]} for s, p in PRODUCTS.items() if q in p["name"].lower()]


@server.tool()
async def convert_from_eur(amount_eur: float, currency: str) -> dict:
    """Convert an amount in EUR to USD, GBP or CHF at today's rate."""
    rate = EUR_RATES.get(currency.strip().upper())
    if rate is None:
        return {"error": f"No rate for {currency}. Supported: USD, GBP, CHF."}
    return {"amount_eur": amount_eur, "currency": currency.upper(), "rate": rate, "amount": round(amount_eur * rate, 2)}


@server.tool()
async def shipping_cost(country: str, weight_kg: float) -> dict:
    """Shipping cost in EUR for a parcel to a country (two-letter code, for example "DE")."""
    fees = SHIPPING.get(country.strip().upper())
    if fees is None:
        return {"error": f"We don't ship to {country}."}
    base, per_kg = fees
    return {"country": country.upper(), "weight_kg": weight_kg, "cost_eur": round(base + per_kg * weight_kg, 2)}


if __name__ == "__main__":
    server.run()

Some answers need one call. Others need the model to chain calls, such as order to customer to shipping, or to notice that search_products returns no stock, so it has to fetch each product.

Step 02

Make the model a setting

This is the whole agent. The model comes from AGENT_MODEL, with gpt-5-mini as the default:

Pythonagent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
import asyncio
import os
import sys

from promptise import build_agent
from promptise.config import StdioServerSpec

# The only line that knows which model runs the agent.
MODEL = os.environ.get("AGENT_MODEL", "openai:gpt-5-mini")


async def main():
    agent = await build_agent(
        model=MODEL,
        servers={"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"])},
        instructions="You are a shop assistant. Look everything up with your tools; never guess.",
        trace_tools=True,
    )
    try:
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Where is order A-1001, and who ordered it?"}]}
        )
        print(f">>> [{MODEL}]", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())

Run it with the default, then with another model:

>_Terminal
python agent.py
AGENT_MODEL=openai:gpt-4.1-mini python agent.py
Output
→ Invoking tool: get_order with {'order_id': 'A-1001'}
✔ Tool result from get_order: {"order_id": "A-1001", "customer_id": "C-8", "items": [{"sku": "MUG-01", "qty": 2}], "status": "shipped", "total_eur": 25.0}
→ Invoking tool: get_customer with {'customer_id': 'C-8'}
✔ Tool result from get_customer: {"customer_id": "C-8", "name": "Lee Park", "country": "DE", "tier": "standard"}
>>> [openai:gpt-5-mini] Order A-1001 is marked as "shipped" in our system — there is no tracking/location info available here. It was ordered by Lee Park (customer ID C-8), country: DE (Germany). 
…
→ Invoking tool: get_order with {'order_id': 'A-1001'}
✔ Tool result from get_order: {"order_id": "A-1001", "customer_id": "C-8", "items": [{"sku": "MUG-01", "qty": 2}], "status": "shipped", "total_eur": 25.0}
→ Invoking tool: get_customer with {'customer_id': 'C-8'}
✔ Tool result from get_customer: {"customer_id": "C-8", "name": "Lee Park", "country": "DE", "tier": "standard"}
>>> [openai:gpt-4.1-mini] Order A-1001 was placed by Lee Park from Germany (country code DE). The order status is "shipped." Is there anything else you would like to know about this order?

Both models made the same two calls with the same arguments. The replies differ in length and tone, which is normal: each model writes its own words. That's also why you can't judge models by reading a few replies. You need tasks with answers you can check.

Step 03

Write tasks with known answers

An eval is a list of questions, the answer each one should get, and a function that checks it. Ask the agent to end with a fixed Answer: line so a few lines of code can grade it. This is the complete script:

Pythoneval_models.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
import asyncio
import json
import re
import statistics
import sys
import time

from promptise import build_agent
from promptise.config import StdioServerSpec

MODELS = sys.argv[1:] or ["openai:gpt-5-mini"]
REPEATS = 3

INSTRUCTIONS = (
    "You are a shop assistant. Look everything up with your tools; never guess. "
    "End your reply with one line: 'Answer: <value>', where the value is only a number "
    "(no currency sign), a SKU, a two-letter country code or a single word."
)

# Ten tasks with one right answer each: single lookups, a "we can't", chains of
# 2-4 calls, parallel calls, filtering and a little arithmetic.
TASKS = [
    ("What's the status of order A-1005?", "shipped"),
    ("Can we ship a 1 kg parcel to Japan? Answer yes or no.", "no"),
    ("Which country does the customer behind order A-1002 live in?", "PT"),
    ("What is the total of order A-1003 in US dollars?", 542.41),
    ("How many orders has customer C-7 placed, including cancelled ones?", 3),
    ("Which is cheaper per unit, MUG-01 or CUP-02? Answer with the SKU.", "CUP-02"),
    ("How much does it cost to ship a 2.5 kg parcel to the customer of order A-1001?", 7.90),
    ("What is the combined stock of every lamp we sell?", 19),
    ("What is the total value in EUR of customer C-7's orders that were not cancelled?", 538.00),
    ("Customer C-9 wants one floor lamp shipped to them; the parcel weighs 4 kg. "
     "What do they pay in Swiss francs for the lamp plus shipping?", 118.06),
]


def grade(reply: str, expected) -> tuple[bool, str]:
    """Pull the 'Answer:' line out of the reply and compare it with the known answer."""
    found = re.findall(r"answer:\s*(.+)", reply, flags=re.IGNORECASE)
    if not found:
        return False, "(no Answer line)"
    value = found[-1].strip().strip("*`. ")
    if isinstance(expected, str):
        return value.upper() == expected.upper(), value
    number = re.search(r"-?\d[\d,]*\.?\d*", value)
    if number is None:
        return False, value
    return abs(float(number.group().replace(",", "")) - expected) < 0.006, value


async def run_task(agent, prompt: str) -> dict:
    start = time.perf_counter()
    result = await agent.ainvoke({"messages": [{"role": "user", "content": prompt}]})
    seconds = time.perf_counter() - start
    messages = result["messages"]
    usage = [m.usage_metadata for m in messages if getattr(m, "usage_metadata", None)]
    tools = [m for m in messages if getattr(m, "type", "") == "tool"]
    return {
        "seconds": seconds,
        "input_tokens": sum(u["input_tokens"] for u in usage),
        "output_tokens": sum(u["output_tokens"] for u in usage),
        "tool_calls": len(tools),
        "reply": messages[-1].content,
    }


async def evaluate(model: str, log) -> dict:
    agent = await build_agent(
        model=model,
        servers={"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"])},
        instructions=INSTRUCTIONS,
    )
    runs = []
    try:
        for repeat in range(REPEATS):
            for number, (prompt, expected) in enumerate(TASKS, 1):
                run = await run_task(agent, prompt)
                run["correct"], run["answer"] = grade(run["reply"], expected)
                run.update(model=model, task=number, repeat=repeat, expected=expected)
                log.write(json.dumps(run) + "\n")
                runs.append(run)
                mark = "ok  " if run["correct"] else "MISS"
                print(f"{model} task {number:2} run {repeat + 1}: {mark} {run['answer']!r} (expected {expected!r})", flush=True)
    finally:
        await agent.shutdown()
    return {
        "model": model,
        "correct": sum(r["correct"] for r in runs),
        "runs": len(runs),
        "median_s": statistics.median(r["seconds"] for r in runs),
        "input_tokens": statistics.mean(r["input_tokens"] for r in runs),
        "output_tokens": statistics.mean(r["output_tokens"] for r in runs),
        "tool_calls": statistics.mean(r["tool_calls"] for r in runs),
    }


async def main():
    with open("eval_runs.jsonl", "a") as log:
        summaries = [await evaluate(model, log) for model in MODELS]
    print(f"\n{'model':<22}{'correct':>10}{'median s':>10}{'in tok':>9}{'out tok':>9}{'tools':>7}")
    for s in summaries:
        print(
            f"{s['model']:<22}{s['correct']:>4}/{s['runs']:<5}{s['median_s']:>10.1f}"
            f"{s['input_tokens']:>9.0f}{s['output_tokens']:>9.0f}{s['tool_calls']:>7.1f}"
        )


asyncio.run(main())

What to notice:

  • The model is the only variable. evaluate builds the same agent, with the same server and instructions, for each model string you pass.

  • Tokens come from the model's own usage data. Each model reply in the result carries the provider's usage_metadata, so the script adds up input and output tokens across all the replies in a task.

  • Every task runs three times. Models aren't deterministic, and one lucky run proves little. Each run goes into eval_runs.jsonl, so you can read the replies behind every miss.

  • The tasks mix difficulty on purpose. The last one asks for a price "in Swiss francs" for a customer in the US. A careless model reads the currency as the destination.

Step 04

Run it across models

Pass the models you want to compare:

>_Terminal
python eval_models.py openai:gpt-5-mini openai:gpt-5 openai:gpt-5-nano openai:gpt-4.1-mini openai:gpt-4.1-nano

That's 150 agent runs. Here's the start, two of the misses, and the summary it printed:

Output
openai:gpt-5-mini task  1 run 1: ok   'shipped' (expected 'shipped')
openai:gpt-5-mini task  2 run 1: ok   'no' (expected 'no')
openai:gpt-5-mini task  3 run 1: ok   'PT' (expected 'PT')
…
openai:gpt-5-nano task  7 run 1: MISS 'DE' (expected 7.9)
…
openai:gpt-4.1-nano task 10 run 3: MISS '0' (expected 118.06)

model                    correct  median s   in tok  out tok  tools
openai:gpt-5-mini       30/30          4.9     2121      472    2.4
openai:gpt-5            30/30          5.0     1974      548    2.4
openai:gpt-5-nano       28/30          6.0     1994     1139    2.4
openai:gpt-4.1-mini     30/30          2.5     1741       84    2.4
openai:gpt-4.1-nano     25/30          2.1     1877       84    2.7

Token counts are averages per task. A small script over eval_runs.jsonl added the slowest single task for each model. Together:

Model

Correct

Median latency

Slowest task

Input tokens per task

Output tokens per task

gpt-5-mini

30/30

4.9 s

10.3 s

2,121

472

gpt-5

30/30

5.0 s

23.0 s

1,974

548

gpt-5-nano

28/30

6.0 s

22.1 s

1,994

1,139

gpt-4.1-mini

30/30

2.5 s

4.6 s

1,741

84

gpt-4.1-nano

25/30

2.1 s

3.7 s

1,877

84

These numbers are from one machine on one day, against one small toolset. Treat them as a result about these tasks, not a ranking of models.

Step 05

Read the misses, not just the score

The score tells you which model to look at. The replies tell you why. These are the seven misses from eval_runs.jsonl, grouped:

  • Right facts, wrong shape. Three of gpt-4.1-nano's misses were correct answers without the Answer: line, such as "The order A-1005 has been shipped." gpt-5-nano once wrote "Shipping cost to DE for a 2.5 kg parcel: 7.9 EUR. Answer: DE". If your code parses the model's output, format slips are real failures.

  • Misreading the task. In one run gpt-4.1-nano replied "there is no shipping available to Switzerland (CH)" for the US customer who wanted to pay in Swiss francs, and answered 0. gpt-5-nano answered "unshippable" to the same task.

  • Arithmetic. gpt-4.1-nano once converted the 37 EUR shipping fee to 34.78 CHF. At the shop's rate it's 34.67, so its total was 11 cents off.

And the reasons to look past accuracy:

  • Three models tied at the top. When several models score 30/30, the eval is too easy to separate them. Add the questions your agent gets wrong in production until the scores spread out.

  • Cheap per token isn't cheap per task. gpt-5-nano used about 13 times the output tokens of gpt-4.1-nano on the same tasks. The GPT-5 models spend output tokens on reasoning before they answer; the usage data reports them as reasoning. Multiply the token columns by your provider's current prices before you decide.

  • Latency has a tail. gpt-5's median was 5 seconds, but one task took 23. For a chat UI, the slow tail matters as much as the middle.

On these tasks, gpt-4.1-mini matched the larger models on accuracy at about half the latency and a fraction of the output tokens. On your tasks the answer may be different, which is the point of running them.


[05]

Switch providers without rewriting the agent

You've seen the string form. Before you switch, check that the new provider is set up, then pick the form that suits where your credentials live.

Check a model string before you use it

promptise models check tells you what a string resolves to and what's missing, without calling anything. --ping adds one real call. On this machine only OpenAI is configured:

>_Terminal
promptise models check openai:gpt-5-mini --ping
promptise models check anthropic:claude-sonnet-4-5
Output
openai:gpt-5-mini → openai:gpt-5-mini  (OpenAI)
  model part: gpt-5-mini — the model name
  route: native integration (core)
  OPENAI_API_KEY: set — platform.openai.com → API keys
…
Usable.
Pinging… ok — replied 'ok'
anthropic:claude-sonnet-4-5 → anthropic:claude-sonnet-4-5  (Anthropic Claude)
  model part: claude-sonnet-4-5 — the model name
  route: native integration (core)
  ANTHROPIC_API_KEY: MISSING — console.anthropic.com → API keys
Not usable yet.
  - ANTHROPIC_API_KEY is not set — console.anthropic.com → API keys (e.g. sk-ant-...)
…

Azure needs three values, and the check names each one and where to find it in the portal:

>_Terminal
promptise models check azure:chat-prod
Output
azure:chat-prod → azure_openai:chat-prod  (Azure OpenAI (OpenAI models deployed in Azure AI Foundry))
  model part: chat-prod — in the string form, your DEPLOYMENT name (Foundry → Deployments → Name); with Model(...), the model name (gpt-4o) — the deployment goes in deployment=
  route: native integration (core)
  AZURE_OPENAI_ENDPOINT: MISSING — Azure AI Foundry portal → your resource → Overview → Endpoint (https://<resource>.openai.azure.com/, no path)
  AZURE_OPENAI_API_KEY: MISSING — Azure AI Foundry portal → your resource → Keys and Endpoint → KEY 1 (or Entra ID: pass extra={'azure_ad_token_provider': ...})
  OPENAI_API_VERSION: MISSING — the REST API version your deployment supports, e.g. 2024-10-21 (Azure docs → 'API version lifecycle')
…
Not usable yet.
…

Note the first line: in the string form, the part after azure: is your deployment name, not the model name. Azure routes requests by deployment.

promptise models env prints the lines to put in your .env file. Promptise loads a .env from the working directory, or a parent up to the project root, and a variable you've already exported wins over the file:

>_Terminal
promptise models env bedrock
Output
# Amazon Bedrock — model string example: bedrock:anthropic.claude-sonnet-4-20250514-v1:0
export AWS_DEFAULT_REGION=us-east-1
#   ↳ the region your Bedrock models are enabled in
export AWS_BEARER_TOKEN_BEDROCK=...
#   ↳ AWS console → Amazon Bedrock → API keys → Generate (a Bedrock API key; long-term keys start with ABSK)

If you switch anyway, build_agent stops before the agent runs, with the same advice:

Output
promptise.models.ModelSetupError: Cannot use model anthropic:'claude-sonnet-4-5' (Anthropic Claude) yet:
  - ANTHROPIC_API_KEY is not set — console.anthropic.com → API keys (e.g. sk-ant-...)
…
  Diagnose any model string with: promptise models check anthropic:claude-sonnet-4-5

To fail a deploy instead of a request, run the same check in CI with check_model. It reads the configuration and calls nothing:

Pythonpreflight.py
1
2
3
4
5
6
7
8
9
10
11
12
import os
import sys

from promptise import check_model

# Check the configuration before deploying. No model is called.
spec = os.environ.get("AGENT_MODEL", "openai:gpt-5-mini")
result = check_model(spec)
print(f"{spec} -> {result.canonical}, ok={result.ok}")
for problem in result.problems:
    print("  -", problem)
sys.exit(0 if result.ok else 1)
Output
openai:gpt-5-mini -> openai:gpt-5-mini, ok=True
bedrock:anthropic.claude-sonnet-4-20250514-v1:0 -> bedrock:anthropic.claude-sonnet-4-20250514-v1:0, ok=False
  - AWS_DEFAULT_REGION is not set — the region your Bedrock models are enabled in (e.g. us-east-1)
  - AWS_BEARER_TOKEN_BEDROCK is not set — AWS console → Amazon Bedrock → API keys → Generate (a Bedrock API key; long-term keys start with ABSK)
…

Configure providers in code with Model

When credentials come from a secret manager rather than the environment, pass a Model instead of a string. It uses the same words for every provider: provider, api_key, endpoint, deployment, api_version, region. This script configures six providers and prints which client each one gets and where it would send requests. resolve() builds the client without calling it, so the placeholder keys never leave the machine:

Pythonproviders_in_code.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
"""Configure six providers in code and show where each one would send requests.

Nothing here calls a model: resolve() only builds the client. Only an OpenAI key
exists on this machine, so the other credentials are placeholders.
"""
import os

from promptise import Model

MODELS = [
    Model("gpt-5-mini", provider="openai", api_key=os.environ["OPENAI_API_KEY"]),
    Model("claude-sonnet-4-5", provider="anthropic", api_key="sk-ant-placeholder"),
    Model(
        "gpt-4o",
        provider="azure",
        deployment="chat-prod",
        endpoint="https://my-resource.openai.azure.com/",
        api_key="placeholder",
        api_version="2024-10-21",
    ),
    Model("anthropic.claude-sonnet-4-20250514-v1:0", provider="bedrock", region="eu-central-1", api_key="ABSK-placeholder"),
    Model("gemini-2.5-pro", provider="gemini", api_key="AIza-placeholder"),
    Model("llama3.1", provider="ollama"),
]

for model in MODELS:
    llm = model.resolve()
    url = (
        getattr(llm, "azure_endpoint", None)
        or getattr(llm, "openai_api_base", None)
        or getattr(llm, "anthropic_api_url", None)
        or "(the OpenAI default)"
    )
    print(f"{model.spec:<55} {type(llm).__name__:<16} {url}")

print(repr(MODELS[2]))
Output
openai:gpt-5-mini                                       ChatOpenAI       (the OpenAI default)
anthropic:claude-sonnet-4-5                             ChatAnthropic    https://api.anthropic.com
azure_openai:gpt-4o                                     AzureChatOpenAI  https://my-resource.openai.azure.com/
bedrock:anthropic.claude-sonnet-4-20250514-v1:0         ChatOpenAI       https://bedrock-runtime.eu-central-1.amazonaws.com/openai/v1
google_genai:gemini-2.5-pro                             ChatOpenAI       https://generativelanguage.googleapis.com/v1beta/openai/
ollama:llama3.1                                         ChatOpenAI       http://localhost:11434/v1
Model(model='gpt-4o', provider='azure', deployment='chat-prod', endpoint='https://my-resource.openai.azure.com/', api_version='2024-10-21', region=None, project=None, temperature=None, max_tokens=None, timeout=None, native=False)

Three things to read from that:

  • Bedrock, Gemini and Ollama all go through `ChatOpenAI`, pointed at each provider's OpenAI-compatible URL. The region you pass for Bedrock lands in the URL.

  • Azure gets its own client with your resource endpoint, and deployment names the deployment the request goes to.

  • `repr()` leaves the key out, so a Model is safe to log. Promptise's source keeps api_key and extra out of repr() and str() on purpose.

Hand any of these to build_agent(model=...) exactly where the string went. To use Bedrock with IAM or SSO credentials instead of a Bedrock API key, install langchain-aws and pass native=True.


[06]

Use any OpenAI-compatible API from Python

vLLM, LM Studio, llama.cpp's server, LiteLLM and most corporate gateways speak the OpenAI chat API. To use one, keep the openai provider and point endpoint at the server's /v1 URL.

To prove the wiring without a GPU, this guide uses a stub: under fifty lines that answer /v1/chat/completions, ask for one tool call, then echo the tool's result. It is not a model, and it only shows that requests, tools and results flow correctly.

Pythonfake_openai_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
"""A stub that speaks just enough of the OpenAI chat API to test the wiring.

It is not a model. It asks for one tool call, then echoes the tool's result.
"""
import json
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer


class ChatCompletions(BaseHTTPRequestHandler):
    def do_POST(self):
        body = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
        tools = [t["function"]["name"] for t in body.get("tools", [])]
        last = body["messages"][-1]
        print(f"POST {self.path} model={body['model']} tools={tools} last={last['role']}", flush=True)

        if self.path != "/v1/chat/completions":
            self.send_error(404)
            return
        if last["role"] == "tool":
            message = {"role": "assistant", "content": f"Stub reply. The tool returned: {last['content']}"}
            finish = "stop"
        else:
            call = {"id": "call_1", "type": "function",
                    "function": {"name": "get_order", "arguments": json.dumps({"order_id": "A-1001"})}}
            message = {"role": "assistant", "content": None, "tool_calls": [call]}
            finish = "tool_calls"

        reply = json.dumps({
            "id": "chatcmpl-stub", "object": "chat.completion", "created": int(time.time()),
            "model": body["model"],
            "choices": [{"index": 0, "message": message, "finish_reason": finish}],
            "usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0},
        }).encode()
        self.send_response(200)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(reply)))
        self.end_headers()
        self.wfile.write(reply)

    def log_message(self, *args):
        pass  # one line per request is printed above


if __name__ == "__main__":
    ThreadingHTTPServer(("127.0.0.1", 8400), ChatCompletions).serve_forever()

With the stub running, the agent from step 2 works against it without a code change. Set the endpoint with OPENAI_BASE_URL, and give the gateway its own key so your real OpenAI key isn't sent to it:

>_Terminal
OPENAI_BASE_URL=http://127.0.0.1:8400/v1 OPENAI_API_KEY=local AGENT_MODEL=openai:stub-model python agent.py
Output
→ Invoking tool: get_order with {'order_id': 'A-1001'}
✔ Tool result from get_order: {"order_id": "A-1001", "customer_id": "C-8", "items": [{"sku": "MUG-01", "qty": 2}], "status": "shipped", "total_eur": 25.0}
>>> [openai:stub-model] Stub reply. The tool returned: {"order_id": "A-1001", "customer_id": "C-8", "items": [{"sku": "MUG-01", "qty": 2}], "status": "shipped", "total_eur": 25.0}

And the stub's log of what it received:

Output
POST /v1/chat/completions model=stub-model tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=user
POST /v1/chat/completions model=stub-model tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=tool

Two requests: the first carried all seven MCP tools in OpenAI's format, and the second carried the tool's result back. That round trip is everything an agent needs from an endpoint. Whether a real model behind it picks the right tool is what your eval measures.

In code, the same endpoint is Model("stub-model", provider="openai", endpoint="http://127.0.0.1:8400/v1", api_key="local").


[07]

Use Ollama with a Python agent

Ollama serves local models on http://localhost:11434, and Promptise reaches it through Ollama's OpenAI-compatible /v1 route. No key is needed. In code, it's Model("llama3.1", provider="ollama"), plus endpoint= when Ollama runs on another machine.

Ollama isn't installed on the machine this guide was written on, which shows a limit of check worth knowing. The configuration is valid, so it says usable. Only --ping finds out nothing is listening:

>_Terminal
promptise models check ollama:llama3.1 --ping
Output
ollama:llama3.1 → ollama:llama3.1  (Ollama (local models))
  model part: llama3.1 — a model you have pulled (`ollama pull llama3.1`)
  route: OpenAI-compatible endpoint {endpoint}/v1 (core, nothing to install)
  OLLAMA_HOST: unset (optional) — only if Ollama is not on the default http://localhost:11434
  No API key. The model must support tool calling to drive MCP tools.
Usable.
Pinging… 
OpenAIConnectionError: Connection error.

To test the Ollama route itself, point it at the stub. This script runs the shop agent twice, once through the plain OpenAI-compatible route and once through the Ollama route:

Pythonstub_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
import asyncio
import sys

from promptise import Model, build_agent
from promptise.config import StdioServerSpec

MODELS = {
    # Any server that speaks the OpenAI chat API: vLLM, LM Studio, a gateway, or our stub.
    "OpenAI-compatible": Model("stub-model", provider="openai", endpoint="http://127.0.0.1:8400/v1", api_key="local"),
    # The Ollama route, pointed at the same stub instead of a real Ollama.
    "Ollama route": Model("llama3.1", provider="ollama", endpoint="http://127.0.0.1:8400"),
}


async def main():
    for label, model in MODELS.items():
        agent = await build_agent(
            model=model,
            servers={"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"])},
            trace_tools=True,
        )
        try:
            result = await agent.ainvoke({"messages": [{"role": "user", "content": "Where is order A-1001?"}]})
            print(f">>> {label}: {result['messages'][-1].content}\n")
        finally:
            await agent.shutdown()


asyncio.run(main())

The stub's log shows the Ollama request arriving at /v1/chat/completions with the model name llama3.1. Promptise added the /v1 itself:

Output
POST /v1/chat/completions model=stub-model tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=user
POST /v1/chat/completions model=stub-model tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=tool
POST /v1/chat/completions model=llama3.1 tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=user
POST /v1/chat/completions model=llama3.1 tools=['get_order', 'list_orders', 'get_customer', 'get_product', 'search_products', 'convert_from_eur', 'shipping_cost'] last=tool

With a real Ollama, choose a model that supports tool calling, then run eval_models.py ollama:<model> next to your hosted candidates. Small local models are where the format and arithmetic slips from step 5 tend to show up, so the eval earns its keep there.


[08]

Honest limits

  • Only OpenAI models ran here. Anthropic, Azure OpenAI, Bedrock and Gemini were configured and checked, and their clients were built, but no request reached them. The Ollama and OpenAI-compatible routes were tested against a stub, not a model.

  • `models check` without `--ping` checks configuration, not reachability. It said Ollama was usable with nothing listening. Use --ping before you rely on a new endpoint.

  • Ten tasks are a small sample. Three runs each smooth out some noise, not all of it. Rerun the eval when a model or your tools change, and grow it from real failures.

  • The grader is strict on format. A correct answer without the Answer: line counts as a miss. That's right if your code parses replies, and too harsh if a person reads them; decide which you need.

  • Tool calling is required. The agent drives MCP tools through the model's function-calling API. A model or server without it doesn't fail loudly: the Promptise docs say the agent loops or replies in prose, with no tool calls in the trace. vLLM, for example, needs --enable-auto-tool-choice and a --tool-call-parser.

  • Tokens aren't dollars. The eval reports tokens. Prices change and differ between providers, so apply your own.


[09]

Frequently asked questions

What is the best LLM for tool calling?

There isn't one answer for every agent. In this guide's eval on OpenAI models, gpt-4.1-mini, gpt-5-mini and gpt-5 all answered 30 of 30 tool-calling tasks correctly, while the nano models missed 2 and 5. Public leaderboards like the Berkeley Function Calling Leaderboard are a good shortlist; your own eval makes the decision.

How do I switch LLM providers in Python without rewriting my agent?

Keep the model out of your code. Pass build_agent a model string such as openai:gpt-5-mini or anthropic:claude-sonnet-4-5 read from configuration, or a Model object with the credentials. Your tools stay on MCP servers, so nothing else changes. Run promptise models check <model> before you switch.

How do I use an OpenAI-compatible API in Python?

Use the openai provider and set the endpoint to the server's /v1 URL: Model("my-model", provider="openai", endpoint="http://host:8000/v1", api_key="..."), or OPENAI_BASE_URL with the string openai:my-model. This works for vLLM, LM Studio, LiteLLM and similar servers, as long as the model behind them supports tool calling.

How do I use Ollama with a Python agent?

Pull a model that supports tool calling, then use ollama:<model> as the model string. Promptise sends requests to http://localhost:11434/v1 by default; set OLLAMA_HOST or endpoint= for another machine. Check it with promptise models check ollama:<model> --ping.

Do I need a different package for each provider?

No. pip install promptise covers OpenAI, Anthropic, Azure OpenAI and every provider reached through an OpenAI-compatible endpoint, including Bedrock, Gemini and Ollama. You only install a provider's own LangChain package when you opt into it with native=True, for example langchain-aws for IAM credentials on Bedrock.


[10]

Where to go next

  • Models & Providers: every provider, every Model word, custom endpoints and troubleshooting.

  • Model Setup: from zero to a working model in two minutes.

  • Best LLMs for Agentic Use Cases: Promptise's own shortlist of models to try.

  • Configuration & Secrets: where keys live and which source wins.

  • Model fallback: try another provider automatically when one fails.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: give the agent your real API to evaluate against.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths