You changed one line in a prompt, and something three features away broke. Nobody noticed for a week. Testing LLM prompts fixes that, but not the way you test normal code: the same input can give a different answer on the next run, and "correct" is sometimes a judgement call. In this guide you'll build a real pytest suite for a prompt that turns support emails into tickets: free unit tests with a mocked model, regression tests against the real model, a second model that grades what code can't, and tests for an AI agent's tool calls. You'll watch it catch real failures, measure how flaky each case is, and see what every run costs. Every snippet ran against Promptise Foundry 1.2.1 and openai:gpt-5-mini, and every output is what it printed.
How do you test LLM prompts?
Treat the prompt as a function with a typed result, and test it in layers. Unit tests mock the model and check your own code: parsing, validation and guards. They're free and run on every push. Regression tests run a fixed set of real inputs with known answers through the real model and assert on the fields, with plain code. For what code can't check, such as whether a summary invents facts, a second model grades the output.
With Promptise Foundry, a prompt is a decorated Python function. Its docstring is the template, and its return type is what you get back:
@prompt(model="openai:gpt-5-mini")
async def extract_ticket(email: str) -> Ticket:
"""You triage email for an online shop's support team.
Read the email and fill in a ticket.
Email:
{email}
"""Because the result is a typed object, a regression test is ordinary pytest:
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
async def test_extracts_the_right_fields(case):
ticket = await ticket_for(case)
for field, expected in case["expect"].items():
assert getattr(ticket, field) == expected, f"{field} was {getattr(ticket, field)!r}"The rest of this guide builds the whole suite around those two pieces, then points the same approach at an agent.
[02]
How the test layers fit together
Each layer answers a different question, at a different price. Run the cheap ones all the time and the expensive ones when they can tell you something new.
Rendering diagram…
Layer | What it catches | Promptise piece | When to run it | Time and cost in this guide |
|---|---|---|---|---|
Unit tests | Parsing, schema, guards, prompt text | PromptTestCase, mock_llm | Every push | 0.64 s, $0 |
Golden cases | Wrong fields from the real model | @prompt with a typed result | When a prompt or model changes | About a minute, about $0.02 |
LLM judge | Summaries that invent facts | A second @prompt | In the same run | Included above |
Repeated runs | Cases that pass only sometimes | registry, a small script | Before you ship a prompt version | About $0.10 to $0.13 for 10 runs |
Agent scenarios | Wrong tool, wrong order, a forbidden call | build_agent, MCP TestClient | When a prompt or tool changes | About 15 s, $0.003 |
[03]
What you need
Python 3.10 or newer.
Promptise Foundry, pytest and pytest-asyncio. This guide used pytest 9.1.1 and pytest-asyncio 1.4.0; pydantic comes with Promptise.
An API key for a model provider. The examples use OpenAI's gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.
pip install promptise pytest pytest-asyncio
export OPENAI_API_KEY="sk-..."[04]
Build a prompt regression suite, step by step
The prompt reads a support email and fills in a ticket: who wrote it, which order, what kind of problem, how urgent, and a one-line summary. It's small enough to read in one go and has the same failure modes as a big one.
Write the prompt as a typed function
@prompt turns a function into a prompt. The docstring is the template, the arguments fill {email}, and the return type tells Promptise what to send back. For a pydantic model, Promptise adds the model's JSON schema to the prompt and parses the answer into a Ticket, so a malformed answer fails loudly instead of flowing on as a dict.
from typing import Literal
from pydantic import BaseModel, Field
from promptise.prompts import prompt
from promptise.prompts.registry import version
class Ticket(BaseModel):
customer_email: str
order_id: str | None = Field(description='The order ID, for example "A-1042", or null if the email has none.')
category: Literal["refund", "shipping", "account", "other"]
urgency: Literal["low", "normal", "high"]
summary: str = Field(description="One sentence a support agent can read at a glance.")
@version("1.0.0")
@prompt(model="openai:gpt-5-mini")
async def extract_ticket(email: str) -> Ticket:
"""You triage email for an online shop's support team.
Read the email and fill in a ticket.
Email:
{email}
"""@version("1.0.0") registers this prompt in Promptise's prompt registry. You'll use it in step 6 to compare versions side by side. To see the exact text the model receives, call extract_ticket.render(email="..."); it builds the prompt, schema included, without calling the model.
Collect real cases with known answers
A regression suite is only as good as its cases. Take them from real traffic, and keep the awkward ones: the customer who types the order ID in lowercase, the email forwarded by a colleague, the one in German. Each case says which fields a correct ticket must have. This guide uses eight; here are three:
CASES = [
{
"id": "broken-item",
"email": "From: dana@example.com\nSubject: Broken mug\n\n"
"Hi, my order A-1001 arrived today and the mug is broken. I'd like my money back please.",
"expect": {"customer_email": "dana@example.com", "order_id": "A-1001", "category": "refund"},
},
{
"id": "lowercase-order-id",
"email": "From: priya@example.com\nSubject: wrong size\n\n"
"hey, got order a1042 yesterday but the shoes are a size too small. can i return them for a refund?",
"expect": {"customer_email": "priya@example.com", "order_id": "A-1042", "category": "refund"},
},
{
"id": "forwarded",
"email": "From: frontdesk@shop.example\nSubject: Fwd: package\n\n"
"Forwarding this from a customer who called in.\n\n"
"---------- Forwarded message ----------\n"
"From: mira@example.com\n"
"My package for order A-1110 says delivered but it's not here.",
"expect": {"customer_email": "mira@example.com", "order_id": "A-1110", "category": "shipping"},
},
]Only assert on fields with one right answer. urgency is left out on purpose: two careful people could disagree about it, and a test that fails on a judgement call teaches you to ignore failures.
Check fields with code and summaries with a judge
The summary is free text, so no assert == can check it. A second model can: give it the email, the summary and a short rubric, and ask for a verdict. That's what people mean by LLM as a judge. Promptise doesn't ship a judge for prompts, but a judge is just another prompt:
from promptise.prompts import prompt
@prompt(model="openai:gpt-5-mini")
async def grade_summary(email: str, summary: str) -> str:
"""You review ticket summaries written for a support team.
A summary is faithful when every fact in it is stated in the email
and it names the customer's main problem. Wording may differ.
Answer PASS or FAIL on the first line, then one sentence explaining why.
Email:
{email}
Summary:
{summary}
"""
def is_faithful(verdict: str) -> bool:
return verdict.strip().upper().startswith("PASS")The verdict is plain text on purpose. The first version of this judge returned a pydantic model. Now and then the model answered with the shape of the JSON schema instead of a verdict, and the parse error crashed a test run. PASS or FAIL on one line can't go wrong that way, and the reason is still there for you to read.
Now the test file. Fields are checked with code. The summary goes to the judge three times, and the majority wins, because the judge is a model too and wobbles. The last test checks the judge itself: given a summary with an invented fact, it has to say FAIL.
import asyncio
import pytest
from cases import CASES
from judge import grade_summary, is_faithful
from ticket_prompt import extract_ticket
pytestmark = pytest.mark.llm
_tickets = {}
async def ticket_for(case):
"""Extract each case once, so both tests below share one model call."""
if case["id"] not in _tickets:
_tickets[case["id"]] = await extract_ticket(case["email"])
return _tickets[case["id"]]
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
async def test_extracts_the_right_fields(case):
ticket = await ticket_for(case)
for field, expected in case["expect"].items():
assert getattr(ticket, field) == expected, f"{field} was {getattr(ticket, field)!r}"
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
async def test_summary_is_faithful(case):
ticket = await ticket_for(case)
# The judge is a model too. Ask it three times and go with the majority.
verdicts = await asyncio.gather(*(grade_summary(email=case["email"], summary=ticket.summary) for _ in range(3)))
assert sum(is_faithful(v) for v in verdicts) >= 2, f"{ticket.summary!r}: {verdicts}"
async def test_the_judge_catches_an_invented_fact():
verdict = await grade_summary(
email=CASES[0]["email"],
summary="The customer wants to exchange the broken mug for a teapot.",
)
assert not is_faithful(verdict)pytestmark = pytest.mark.llm tags every test in the file. That one marker is what lets CI run the free tests on every push and the paid ones only when they matter.
Configure pytest for real model calls
Two small files make the suite behave. pytest.ini declares the marker and runs async tests automatically:
[pytest]
asyncio_mode = auto
# One event loop for the whole run. The OpenAI client is shared between calls,
# and a new loop per test leaves it bound to a closed one.
asyncio_default_test_loop_scope = session
asyncio_default_fixture_loop_scope = session
markers =
llm: calls a real model, so it is slow and costs moneyThe loop scope matters more than it looks. pytest-asyncio gives each test a fresh event loop by default, while LangChain, which Promptise uses to call models, caches one HTTP client for the whole process. The first test works and the next ones fail with RuntimeError: Event loop is closed. One loop per session avoids it.
conftest.py adds up the tokens of every model call and prints the cost at the end of the run:
import pytest
from langchain_core.callbacks import get_usage_metadata_callback
# USD per million tokens for openai:gpt-5-mini. Check your provider's current prices.
PRICE_INPUT = 0.25
PRICE_OUTPUT = 2.00
totals = {"input": 0, "output": 0}
@pytest.fixture(autouse=True)
def count_tokens():
"""Add up the tokens of every model call a test makes."""
with get_usage_metadata_callback() as usage:
yield
for model_usage in usage.usage_metadata.values():
totals["input"] += model_usage["input_tokens"]
totals["output"] += model_usage["output_tokens"]
def pytest_terminal_summary(terminalreporter):
if totals["input"]:
cost = (totals["input"] * PRICE_INPUT + totals["output"] * PRICE_OUTPUT) / 1_000_000
terminalreporter.write_line(
f"LLM usage: {totals['input']} input + {totals['output']} output tokens = ${cost:.4f}"
)The prices are OpenAI's list prices for gpt-5-mini at the time of writing. Promptise's own per-call stats have input_tokens and output_tokens fields, but in 1.2.1 they stay at 0, which is why the count comes from LangChain's usage callback instead.
Run it and read the failures
pytest -m llmcollected 27 items / 7 deselected / 20 selected
test_agent.py ... [ 15%]
test_ticket_llm.py ...F.........F... [100%]
=================================== FAILURES ===================================
______________ test_extracts_the_right_fields[lowercase-order-id] ______________
…
E AssertionError: order_id was 'a1042'
E assert 'a1042' == 'A-1042'
…
_____________________ test_summary_is_faithful[forwarded] ______________________
…
E AssertionError: 'Customer reports order A-1110 shows as delivered but package is not received; called in and forwarded for investigation.': ['FAIL — The summary correctly restates that order A-1110 is marked delivered but not received and that the customer called in, but it adds "forwarded for investigation," which is not stated in the email.', …]
E assert 0 >= 2
…
LLM usage: 7816 input + 10140 output tokens = $0.0222
=========================== short test summary info ============================
FAILED test_ticket_llm.py::test_extracts_the_right_fields[lowercase-order-id]
FAILED test_ticket_llm.py::test_summary_is_faithful[forwarded] - AssertionErr...
============ 2 failed, 18 passed, 7 deselected in 69.29s (0:01:09) =============The three test_agent.py tests are the agent scenarios from later in this guide; they share the llm marker. Two real problems came out of the other seventeen:
The order ID came back as the customer typed it. Your order system stores A-1042; a1042 won't match anything. The schema said "for example A-1042", and the model took that as a hint, not a rule.
The summary added a fact. Nothing in the email says the message was forwarded "for investigation". All three judge calls caught it. It's a small invention, and exactly the kind that erodes trust in a summary.
Run it again without changing anything, and you'll often get a different result. In the second run here, only the order ID failed. That isn't a broken test. It's the most important thing to understand about testing LLM prompts.
Measure how often each case passes
One run tells you whether a case can fail. To know how often it fails, run each case several times and count. This script does that for any registered version of the prompt, checking fields and asking the judge once per run:
"""Run every case several times against one prompt version and count the passes."""
import asyncio
import sys
from collections import Counter
from langchain_core.callbacks import get_usage_metadata_callback
import ticket_prompt # noqa: F401 registers extract_ticket 2.1.0
import ticket_prompt_v1 # noqa: F401 registers extract_ticket 1.0.0
import ticket_prompt_v2 # noqa: F401 registers extract_ticket 2.0.0
from cases import CASES
from judge import grade_summary, is_faithful
from promptise.prompts.registry import registry
PRICE_INPUT, PRICE_OUTPUT = 0.25, 2.00 # USD per million tokens for openai:gpt-5-mini
limit = asyncio.Semaphore(8)
async def run_once(extract, case):
async with limit:
try:
ticket = await extract(case["email"])
except Exception as error:
return type(error).__name__, False
wrong = [f"{f}={getattr(ticket, f)!r}" for f, v in case["expect"].items() if getattr(ticket, f) != v]
verdict = await grade_summary(email=case["email"], summary=ticket.summary)
return (", ".join(wrong) or "ok"), is_faithful(verdict)
async def main(ver: str, runs: int):
extract = registry.get("extract_ticket", ver)
with get_usage_metadata_callback() as usage:
results = await asyncio.gather(*(run_once(extract, c) for c in CASES for _ in range(runs)))
print(f"extract_ticket {ver}, {runs} runs per case")
for i, case in enumerate(CASES):
mine = results[i * runs : (i + 1) * runs]
fields = sum(r == "ok" for r, _ in mine)
judge = sum(faithful for _, faithful in mine)
misses = Counter(r for r, _ in mine if r != "ok")
print(f" {case['id']:<20} fields {fields:>2}/{runs} judge {judge:>2}/{runs} {dict(misses) or ''}")
tokens = next(iter(usage.usage_metadata.values()))
cost = (tokens["input_tokens"] * PRICE_INPUT + tokens["output_tokens"] * PRICE_OUTPUT) / 1_000_000
print(f"{tokens['input_tokens']} input + {tokens['output_tokens']} output tokens = ${cost:.4f}, "
f"${cost / runs:.4f} per pass over all {len(CASES)} cases")
asyncio.run(main(sys.argv[1], int(sys.argv[2])))The original prompt from step 1 is kept as ticket_prompt_v1.py, and registry.get("extract_ticket", "1.0.0") fetches it by version. Here it is, ten runs per case:
python repeat.py 1.0.0 10extract_ticket 1.0.0, 10 runs per case
broken-item fields 10/10 judge 10/10
late-delivery fields 10/10 judge 9/10
password-reset fields 10/10 judge 10/10
lowercase-order-id fields 3/10 judge 10/10 {"order_id='A1042'": 5, "order_id='a1042'": 2}
two-orders fields 10/10 judge 10/10
forwarded fields 10/10 judge 10/10
german fields 9/10 judge 9/10 {'ValidationError': 1}
no-order fields 10/10 judge 10/10
38057 input + 44835 output tokens = $0.0992, $0.0099 per pass over all 8 casesNow the picture is clear:
The order ID is a real bug. It passed 3 times in 10, in two different wrong spellings. A single CI run lets it through 30% of the time.
The invented "for investigation" didn't come back. Over ten runs the judge passed every forwarded summary. It was a one-off, worth knowing about but not worth rewriting the prompt for.
One answer didn't parse at all. Every failure like this inspected while writing this guide was the model answering with the JSON schema's structure, its fields wrapped in a "properties" object, instead of a ticket. Promptise asks for JSON in the prompt text rather than through the provider's structured output mode, and gpt-5-mini occasionally gets that wrong.
Ten runs of eight cases cost about ten cents. That's cheap enough to do before every prompt change you ship.
Fix the prompt and compare versions
The first attempt at a fix, version 2.0.0, added a rule for order IDs, a pattern on the field so a wrong ID fails validation, and a summary that also says what the customer wants:
summary: str = Field(description="One sentence: the customer's problem and what they want us to do.")It made things worse:
python repeat.py 2.0.0 10extract_ticket 2.0.0, 10 runs per case
broken-item fields 10/10 judge 10/10
late-delivery fields 7/10 judge 0/10 {'ValidationError': 3}
password-reset fields 9/10 judge 0/10 {'ValidationError': 1}
lowercase-order-id fields 6/10 judge 6/10 {'ValidationError': 4}
two-orders fields 8/10 judge 8/10 {'ValidationError': 2}
forwarded fields 10/10 judge 0/10
german fields 7/10 judge 0/10 {'ValidationError': 3}
no-order fields 4/10 judge 4/10 {'ValidationError': 6}
42706 input + 60208 output tokens = $0.1311, $0.0131 per pass over all 8 casesTwo regressions, both caught before anyone shipped them. Asking for "what they want us to do" made the model invent requests. Most of these emails never state one, so it wrote things like "asks us to locate it or provide a replacement or refund", and the judge failed them. And 19 of the 80 answers failed to parse, against 1 in 80 before: the longer prompt made the schema mix-up far more common. A run counts as a judge failure when extraction crashed, which is why those columns drop together.
Version 2.1.0 keeps the order ID rule and the pattern, asks for a summary built only from the email, and says plainly where the fields go. It also adds a guard, which you'll test in the next step:
from typing import Literal
from pydantic import BaseModel, Field
from promptise.prompts import prompt
from promptise.prompts.guards import GuardError, guard, output_validator
from promptise.prompts.registry import version
class Ticket(BaseModel):
customer_email: str
order_id: str | None = Field(
pattern=r"^A-\d{4}$",
description='The order ID, for example "A-1042", or null if the email has none.',
)
category: Literal["refund", "shipping", "account", "other"]
urgency: Literal["low", "normal", "high"]
summary: str = Field(description="One sentence: the customer's problem, using only facts from the email.")
def not_our_own_address(ticket: Ticket) -> Ticket:
"""A forwarded email belongs to the customer, never to our own staff."""
if ticket.customer_email.endswith("@shop.example"):
raise GuardError(f"customer_email is our own address: {ticket.customer_email}", guard_name="own_address")
return ticket
@version("2.1.0")
@prompt(model="openai:gpt-5-mini")
@guard(output_validator(not_our_own_address))
async def extract_ticket(email: str) -> Ticket:
"""You triage email for an online shop's support team.
Read the email and fill in a ticket.
Rules:
- Write order IDs the way our system stores them: a capital A, a hyphen and
four digits, such as A-1234, even when the customer writes a1234 or #A1234.
- In a forwarded email, the customer is the original sender, not our staff.
- Put the ticket's fields at the top level of your JSON, not inside "properties".
Email:
{email}
"""python repeat.py 2.1.0 10extract_ticket 2.1.0, 10 runs per case
broken-item fields 10/10 judge 10/10
late-delivery fields 10/10 judge 7/10
password-reset fields 10/10 judge 10/10
lowercase-order-id fields 10/10 judge 10/10
two-orders fields 10/10 judge 10/10
forwarded fields 10/10 judge 10/10
german fields 10/10 judge 10/10
no-order fields 10/10 judge 10/10
46335 input + 49518 output tokens = $0.1106, $0.0111 per pass over all 8 casesEighty out of eighty on the fields, with no parse errors. The one soft spot is the judge's 7 of 10 on late-delivery. A single judge call is noisy on borderline summaries, which is why the pytest suite asks it three times. If a case stayed low, the judge's reasons would be the next thing to read.
Lock the fix in with free unit tests
Each fix deserves a test that costs nothing to run. PromptTestCase from Promptise patches the model call, so run_prompt returns whatever string you give mock_llm, pushed through the same parsing, validation and guards as a real answer. These tests run in well under a second, with no API key:
import pytest
from pydantic import ValidationError
from promptise.prompts.guards import GuardError
from promptise.prompts.testing import PromptTestCase
from ticket_prompt import Ticket, extract_ticket
GOOD = (
'{"customer_email": "dana@example.com", "order_id": "A-1001", "category": "refund", '
'"urgency": "high", "summary": "The mug in order A-1001 arrived broken; she wants a refund."}'
)
class TestExtractTicket(PromptTestCase):
prompt = extract_ticket
async def test_parses_the_model_json_into_a_ticket(self):
with self.mock_llm(GOOD):
ticket = await self.run_prompt("any email")
self.assert_schema(ticket, Ticket)
assert ticket.order_id == "A-1001"
async def test_accepts_json_inside_a_markdown_fence(self):
with self.mock_llm(f"```json\n{GOOD}\n```"):
ticket = await self.run_prompt("any email")
assert ticket.category == "refund"
async def test_rejects_an_unknown_category(self):
with self.mock_llm(GOOD.replace('"refund"', '"complaint"')):
with pytest.raises(ValidationError):
await self.run_prompt("any email")
async def test_rejects_a_malformed_order_id(self):
with self.mock_llm(GOOD.replace("A-1001", "a1001")):
with pytest.raises(ValidationError):
await self.run_prompt("any email")
async def test_refuses_our_own_address_as_the_customer(self):
with self.mock_llm(GOOD.replace("dana@example.com", "frontdesk@shop.example")):
with pytest.raises(GuardError) as error:
await self.run_prompt("any email")
assert error.value.guard_name == "own_address"
def test_the_prompt_sends_the_schema(self):
text = extract_ticket.render(email="any email")
assert '"required"' in text and "customer_email" in textpytest -m "not llm"collected 27 items / 20 deselected / 7 selected
test_agent.py . [ 14%]
test_ticket_unit.py ...... [100%]
======================= 7 passed, 20 deselected in 0.64s =======================What each kind of test protects:
The `pattern` turns a silent bug into a loud one. Against version 1.0.0, test_rejects_a_malformed_order_id failed with DID NOT RAISE: a wrong ID would have flowed into your system as data. Now it raises, and your code can retry or flag the ticket.
The guard runs on every call, in production too. output_validator wraps any function as an output guard. It runs after the answer is parsed, so it receives the Ticket, and a GuardError stops the bad ticket before your code sees it. Use guards for rules a schema can't express, like "never file a ticket under our own address".
The prompt text is tested too. render() returns the exact text the model gets, so you can assert that a rule or the schema is still there after someone edits the docstring.
If several prompts share the same rules, a PromptSuite applies its default_guards and default_constraints to every prompt in the suite, so one guard covers them all.
These tests can't tell you whether the model follows the rules. That's what the llm suite and repeat.py are for.
[05]
Evaluate an AI agent's tool calls
An agent adds a new kind of failure: it can call the wrong tool, call the right one at the wrong time, or call one it should never touch. You test that by giving it a task and asserting on the tool calls it made, not on its wording.
Here is a two-tool MCP server for the agent:
from promptise.mcp.server import MCPServer, ToolError
server = MCPServer("shop")
# A stand-in for your order system.
ORDERS = {
"A-1001": {"item": "Ceramic mug", "total": 18.50, "status": "delivered"},
"A-1002": {"item": "Standing desk", "total": 1240.00, "status": "in transit"},
}
@server.tool()
async def get_order(order_id: str) -> dict:
"""Look up an order: item, total in USD and delivery status."""
order = ORDERS.get(order_id)
if order is None:
raise ToolError(f"There is no order {order_id}.")
return {"order_id": order_id, **order}
@server.tool()
async def issue_refund(order_id: str, amount: float) -> dict:
"""Refund money to the customer. Only for delivered orders."""
order = ORDERS.get(order_id)
if order is None:
raise ToolError(f"There is no order {order_id}.")
if amount > order["total"]:
raise ToolError(f"Refund {amount} is more than the order total {order['total']}.")
return {"order_id": order_id, "amount": amount, "status": "refunded"}
if __name__ == "__main__":
server.run()The tests come in two kinds. The first checks the tool's own rule with Promptise's MCP TestClient, which calls the tool in memory with no model involved, so it's free and exact. The others run the real agent through three scenarios and check which tools it called, and in what order:
import sys
import pytest
from promptise import build_agent
from promptise.config import StdioServerSpec
from promptise.mcp.server import TestClient
from shop_server import server
SCENARIOS = [
{
"id": "refund-a-delivered-order",
"ask": "Order A-1001 arrived broken. Please refund it in full.",
"must_call": ["get_order", "issue_refund"],
"must_not_call": [],
},
{
"id": "refuse-an-order-in-transit",
"ask": "Refund order A-1002 please, it's late.",
"must_call": ["get_order"],
"must_not_call": ["issue_refund"],
},
{
"id": "answer-without-tools",
"ask": "What are your opening hours?",
"must_call": [],
"must_not_call": ["get_order", "issue_refund"],
},
]
def tool_calls(result):
"""Every tool call the model asked for, in order, as (name, arguments)."""
return [
(call["name"], call["args"])
for message in result["messages"]
for call in getattr(message, "tool_calls", None) or []
]
async def test_the_refund_tool_refuses_more_than_the_total():
"""The tool's own rules, tested without a model."""
result = await TestClient(server).call_tool("issue_refund", {"order_id": "A-1001", "amount": 99.0})
assert "more than the order total" in result[0].text
@pytest.fixture(scope="module")
async def agent():
agent = await build_agent(
model="openai:gpt-5-mini",
servers={"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"])},
instructions=(
"You are a support assistant for an online shop. Look up an order before "
"you act on it. Only refund orders that have been delivered."
),
)
yield agent
await agent.shutdown()
@pytest.mark.llm
@pytest.mark.parametrize("scenario", SCENARIOS, ids=[s["id"] for s in SCENARIOS])
async def test_the_agent_uses_the_right_tools(agent, scenario):
result = await agent.ainvoke({"messages": [{"role": "user", "content": scenario["ask"]}]})
names = [name for name, _ in tool_calls(result)]
for name in scenario["must_call"]:
assert name in names, f"expected a call to {name}, got {names}"
for name in scenario["must_not_call"]:
assert name not in names, f"{name} must not be called, got {names}"
if "issue_refund" in names:
assert names.index("get_order") < names.index("issue_refund"), "refunded before looking up"pytest -m llm test_agent.py -q... [100%]
LLM usage: 1818 input + 1517 output tokens = $0.0035
3 passed, 1 deselected in 17.21sA few things make agent tests hold up:
Assert on actions, not words. tool_calls reads the calls the model asked for from the agent's messages. Whether it says "I've refunded" or "Your refund is on its way" doesn't matter; whether issue_refund ran does.
Test the refusals. The most valuable scenario is the one where the agent must not act: the desk is still in transit, so no refund.
Build the agent once per module. Starting the agent and its MCP server takes a moment, so the fixture shares one agent across the scenarios. Each ainvoke gets only the messages you pass it.
These three scenarios passed in all seven runs made for this guide, 21 out of 21, at about a third of a cent per run. If your tools come from an existing REST API, MCPcast's --eval option goes further: a model writes realistic tasks for the generated server and a real agent tries each one, as shown in OpenAPI to MCP: Turn Any REST API into an MCP Server.
[06]
Run prompt regression tests in CI
Split the suite by price, using the marker:
pytest -m "not llm"
pytest -m llmEvery push: the first command. No API key, no cost, under a second for this suite. It catches broken parsing, a deleted rule, a guard that stopped firing.
When a prompt, a model or a tool changes, and nightly: the second. It needs OPENAI_API_KEY as a CI secret, takes about a minute here and costs about two cents. Running it nightly catches the changes you didn't make: a provider updating a model behind the same name.
Before you ship a new prompt version: repeat.py against the old and new versions, and compare the tables.
Treat infrastructure errors as their own category. While this guide was written, one run failed on a DNS outage and another when the API account ran out of credit, with Error code: 429 and insufficient_quota. Neither says anything about your prompt, so don't let them count as a prompt failure, and don't let a retry policy hide a real one.
[07]
Flakiness and cost: the honest numbers
Everything here is measured on this guide's runs with gpt-5-mini:
What | Time | Cost |
|---|---|---|
Unit tests, 7 tests | 0.64 s | $0 |
pytest -m llm, 20 tests | 59 to 69 s | $0.019 to $0.022 |
Agent scenarios, 3 tests | 14 to 17 s | $0.0030 to $0.0035 |
repeat.py, 10 runs of 8 cases plus the judge | about a minute | $0.10 to $0.13 |
Output tokens dominate. A ticket is a few dozen words, yet output tokens outnumbered input tokens in most runs. gpt-5-mini reasons before it answers, and OpenAI bills reasoning tokens as output, at eight times the input price.
You can't switch the randomness off. The usual advice is temperature 0. For gpt-5-mini that isn't available: LangChain drops any temperature other than the default for GPT-5 reasoning models, and @prompt takes only a model string anyway. Design the suite for variation instead of hoping it away.
Sample size is arithmetic. A case that fails 30% of the time slips through a single run 70% of the time. To see a failure that happens one time in ten with 95% confidence, you need 29 runs, because 0.9 to the power of 29 is about 0.05. Ten runs per case won't prove a prompt is perfect; it will show you which cases need attention.
Retries belong in production, not in your measurements. Promptise's retry() wraps a prompt and tries again on any exception, with exponential backoff that starts at one second, so a ValidationError or GuardError in production gets a second chance. Measure without it: a prompt that needs retries to pass is a prompt you're paying for two or three times.
Judges wobble too. In the final measurement the judge passed late-delivery 7 times in 10, while every field was right. Asking three times and taking the majority costs three judge calls per case, and makes a single unlucky verdict much less likely to fail the run. Keep a test that proves the judge can fail, like test_the_judge_catches_an_invented_fact, so you know it isn't passing everything.
[08]
Honest limits
These are true of Promptise Foundry 1.2.1:
No built-in judge or eval runner for prompts. Promptise gives you typed prompts, PromptTestCase, guards and a version registry. The judge, the cases and the repeat loop are yours, as built above. MCPcast's evaluation covers generated MCP servers only.
Structured output is asked for in the prompt. Promptise appends the JSON schema to the prompt text and parses the reply; it doesn't use the provider's structured output mode. Expect the occasional reply in the schema's shape, keep the "top level" instruction, and catch ValidationError where it matters.
Prompt stats don't count tokens. last_stats has latency_ms, but input_tokens and output_tokens are always 0. Use LangChain's usage callback, as in conftest.py, or agent observability for agents.
The prompt inspector records less than the docs say. PromptInspector only records a trace for prompts built with blocks, and then only the assembled block text: output_text, latency_ms and the guard lists stay empty. To see what a prompt sends, use render().
Some documented options don't exist yet. @prompt(observe=True) is accepted but records nothing, @prompt(observer=...) raises TypeError, and schema_strict(max_retries=2) raises TypeError too. schema_strict() with no arguments works, and only checks string outputs.
The registry is per process. @version registers into an in-memory registry when the module is imported, and registering the same name and version twice raises ValueError. It's a way to load and compare versions in code, not a store that survives a restart.
[09]
Frequently asked questions
How do you evaluate LLM output in Python?
Split it into what code can check and what it can't. Parse the output into a typed object, then assert on the fields that have one right answer, such as IDs, categories and required fields. For free text, use a second model with a short rubric as a judge, ask it more than once, and check now and then that it can fail.
Can you trust an LLM as a judge?
Partly. In this guide the judge caught real invented facts, but it also failed some summaries of an email whose other summaries it passed. Keep the rubric to one or two checkable rules, return a simple PASS or FAIL, take a majority vote, and read its reasons when it fails instead of trusting the count.
How do you unit test an AI agent?
Test its tools without a model, using an in-memory client like Promptise's MCP TestClient. Then run the agent through short scenarios and assert on the tool calls it made: which tools, in what order, and which ones it must never call. Don't assert on its wording.
Should I set temperature to 0 for LLM tests?
When the model supports it, it reduces variation but doesn't remove it, and reasoning models such as gpt-5-mini only accept the default temperature. Build the suite so it holds up under variation: assert on structured fields, run each case several times before shipping, and track pass rates instead of single results.
How much does it cost to test LLM prompts?
Less than people expect. This guide's regression suite costs about two cents a run, and ten runs of every case about a dime, with gpt-5-mini. The unit tests cost nothing, so run them on every push and the paid tests when a prompt or model changes.
[10]
Where to go next
Prompt Testing: PromptTestCase, mock_llm and every assertion.
Guards: built-in guards and how to write your own.
Suite & Registry: group prompts with shared guards, and version them.
Inspector & Observability: tracing prompt assembly, with the limits above in mind.
Prompt chaining: retry, fallback and other ways to compose prompts.
MCP server testing: more on TestClient.
How to Connect MCP Servers to Your AI Agent in Python: the agent and MCP basics behind the agent tests.