An agent gives a customer a wrong answer, and you have to explain why. Was it the prompt, the model, or one of its tools? Without a trace you're guessing. AI agent observability means recording every LLM call and every tool call, with its arguments, result, tokens and latency, so you can replay what the agent did and point at the step that went wrong. In this guide you'll add it to a small shop assistant with two MCP tools, read the real trace, find a real bug with it, estimate cost per question, and send the same events to a log file, a webhook and OpenTelemetry. Every snippet ran against Promptise Foundry 1.2.1, and the output is what it printed.
How do you add observability to an AI agent?
With Promptise Foundry, pass observe=True to build_agent. From then on every model call (tokens, latency, model version, which tools it asked for) and every tool call (arguments, result, latency) is recorded on a timeline. agent.get_stats() returns the totals, and an HTML report is written when the agent shuts down:
from promptise import build_agent
from promptise.config import StdioServerSpec
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
},
instructions="You are a shop assistant. Use your tools to answer. Be brief.",
observe=True,
)
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
print(json.dumps(agent.get_stats(), indent=2))
print(agent.generate_report("shop-report.html"))No trace data leaves your machine unless you add a transporter that sends it. The trace lives in memory and in local files, and the code that records it is open source. The rest of this guide shows what the trace contains and how to use it.
[02]
How agent observability works
When observability is on, build_agent attaches a PromptiseCallbackHandler to every run. It turns each model call and tool call into a timeline event and hands it to an ObservabilityCollector. The collector keeps the events in memory and passes each one to the transporters you chose.
Rendering diagram…
These are the events you'll see from an agent with MCP tools, and what each one carries:
Event | Recorded when | What it carries |
|---|---|---|
llm.start | A model call begins | Model name |
llm.end | A model call returns | Prompt, completion and total tokens, latency, exact model version, the tools the model asked for |
tool.call | A tool starts | Tool name and arguments |
tool.result | A tool returns | The first 2,000 characters of the result, latency |
llm.error | A model call raises | Error type, message and traceback |
Each event can also carry a user_id and session_id, taken from the CallerContext you pass to ainvoke. That turns out to be the easiest way to pull one request back out of a busy trace.
[03]
What you need
Python 3.10 or newer.
Promptise Foundry: pip install promptise. Everything up to the OpenTelemetry section needs nothing else.
For OpenTelemetry export:
pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc. The docs suggest pip install "promptise[all]", which works too but also installs much more, including PyTorch.An API key for a model provider. This guide uses OpenAI, so set OPENAI_API_KEY.
[04]
Trace an AI agent, step by step
You'll build a shop assistant that answers questions about products and stock, run three real questions through it, and use the trace to find out why one answer is wrong.
Give the agent two tools
The tools live in a small MCP server. check_stock waits 0.8 seconds on purpose, standing in for a slow warehouse API, so you can see latency in the trace. search_products has a bug, the kind that slips into real code: it only matches the start of the product name, like a LIKE 'query%' database index.
import asyncio
from promptise.mcp.server import MCPServer, ToolError
server = MCPServer("shop")
# A stand-in for your product catalog and warehouse system.
PRODUCTS = {
"KB-200": {"name": "KB-200 mechanical keyboard", "price_eur": 89.0},
"MS-10": {"name": "MS-10 wireless mouse", "price_eur": 24.0},
"MS-30": {"name": "MS-30 ergonomic mouse", "price_eur": 49.0},
}
STOCK = {"KB-200": 14, "MS-10": 0, "MS-30": 6}
@server.tool()
async def search_products(query: str) -> list[dict]:
"""Search the catalog by product name. Returns SKU, name and price in EUR.
Args:
query: Words to look for in the product name, for example "mouse".
"""
q = query.lower()
# Prefix match, like a LIKE 'query%' index in a database.
return [{"sku": sku, **p} for sku, p in PRODUCTS.items() if p["name"].lower().startswith(q)]
@server.tool()
async def check_stock(sku: str) -> dict:
"""Check how many units of a product are in the warehouse.
Args:
sku: The product SKU, for example "KB-200".
"""
await asyncio.sleep(0.8) # the warehouse API is slow
if sku.upper() not in STOCK:
raise ToolError(f"Unknown SKU {sku}.")
return {"sku": sku.upper(), "units": STOCK[sku.upper()]}
if __name__ == "__main__":
server.run()If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python walks through this part.
Turn on observability and run real questions
observe=True is enough to start. For a trace you can come back to, pass an ObservabilityConfig instead. This one keeps the HTML report and adds the JSON transporter, which streams every event to a file as it happens. The CallerContext tags each question's events with its own session_id.
import asyncio
import json
import sys
from promptise import CallerContext, ObservabilityConfig, TransporterType, build_agent
from promptise.config import StdioServerSpec
QUESTIONS = {
"q1": "Is the MS-30 in stock?",
"q2": "How much is the MS-10?",
"q3": "Do you sell an ergonomic mouse?",
}
async def main():
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
},
instructions="You are a shop assistant. Use your tools to answer. Be brief.",
observe=ObservabilityConfig(
session_name="shop-assistant",
transporters=[TransporterType.HTML, TransporterType.JSON],
output_dir="traces",
),
)
try:
for qid, question in QUESTIONS.items():
# Tag every event of this request, so you can pull it out of the trace later.
caller = CallerContext(user_id="customer-42", metadata={"session_id": qid})
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": question}]}, caller=caller
)
print(f"[{qid}] {question}")
print(f" {result['messages'][-1].content}")
print(json.dumps(agent.get_stats(), indent=2))
finally:
await agent.shutdown()
asyncio.run(main())python shop_agent.py[q1] Is the MS-30 in stock?
Yes — we have 6 units of the MS-30 in stock.
[q2] How much is the MS-10?
The MS-10 wireless mouse costs €24.00. Would you like me to check stock or add one to your cart?
[q3] Do you sell an ergonomic mouse?
I don’t see any ergonomic mice in our catalog right now. Would you like me to search for related items (vertical mice, trackballs, or other ergonomic pointing devices) or check for a specific brand/model?
{
"entry_count": 22,
"agent_count": 1,
"total_duration_s": 13.86664605140686,
"total_prompt_tokens": 2032,
"total_completion_tokens": 715,
"total_tokens": 2747,
"llm_call_count": 7,
"tool_call_count": 4,
"error_count": 0,
"retry_count": 0,
"latency_p50_ms": 1180.9399127960205,
"latency_p95_ms": 2771.5067863464355,
"latency_p99_ms": 2771.5067863464355,
…
"events_by_type": {
"llm.start": 7,
"llm.end": 7,
"tool.call": 4,
"tool.result": 4
}
}Three questions took 7 model calls, 4 tool calls and 2,747 tokens. Two details help you read these numbers. The latency percentiles mix model calls and tool calls, because they cover every event that has a duration. And total_duration_s counts from when the agent was built, not just the time spent answering.
The third answer is wrong. The catalog has an ergonomic mouse, the MS-30, and the agent said it didn't. The trace will tell you why.
Open the HTML report
When the agent shut down, the HTML transporter wrote traces/shop-assistant-report-20261010_184106.html. It's one self-contained file, with no external scripts or styles, that you can open in any browser or attach to a ticket. Rendered in a browser (here, headless Chrome), it shows six stat cards, filter buttons and one row per event:
title: Promptise Agent Report
…
rendered stat cards: [('22', 'Total Events'), ('0', 'Total Tokens'), ('0', 'LLM Calls'), ('0', 'Tool Calls'), ('0', 'Errors'), ('0', 'Cache Hits')]
rendered filter buttons: ['All', 'Tool', 'Llm', 'Error', 'Cache']
rendered rows: 22
📌 llm.start LLM call started 18:40:50
📌 llm.end LLM call completed (288 tokens, 1831.7ms) 18:40:52
📌 tool.call Calling tool: check_stock 18:40:52
📌 tool.result Tool completed: check_stock 18:40:53Each row shows the event type, a one-line description and the time. The token count and latency of each model call are in its description.
You can also write a report at any point with agent.generate_report(path). Be aware that it doesn't use your path as the file name. From the quick example at the top:
print(agent.generate_report("shop-report.html"))shop-report.htmlls *.html reports/shop-report-report-20261010_185045.html
reports/:
promptise-report-20261010_185048.htmlIt printed shop-report.html, but the file it wrote has a timestamp added. The second file is the automatic report from observe=True, which goes to ./reports unless you set output_dir.
Find the step where the agent went wrong
The JSON transporter streamed every event to traces/shop-assistant-events.ndjson, one JSON object per line. Each line has the event type, the session_id you set, and a metadata object with the details. This short script turns that file into numbered steps per question, with the token cost of each:
"""Print a saved Promptise trace as numbered steps, one block per request."""
import json
import sys
from collections import defaultdict
# USD per million tokens for openai:gpt-5-mini. Check your provider's current prices.
PRICE_INPUT = 0.25
PRICE_OUTPUT = 2.00
requests = defaultdict(list)
with open(sys.argv[1]) as f:
for line in f:
event = json.loads(line)
requests[event["session_id"]].append(event)
for request_id, events in requests.items():
print(f"== {request_id}")
step, prompt_tokens, completion_tokens = 0, 0, 0
for e in events:
m = e["metadata"]
if e["event_type"] == "llm.end":
step += 1
prompt_tokens += m["prompt_tokens"]
completion_tokens += m["completion_tokens"]
wants = ", ".join(m.get("tool_calls", [])) or "final answer"
print(f"{step:>2}. model {m['latency_ms']:>6.0f} ms {m['total_tokens']:>4} tokens -> {wants}")
elif e["event_type"] == "tool.call":
step += 1
print(f"{step:>2}. call {m['tool_name']} {m['arguments']}")
elif e["event_type"] == "tool.result":
step += 1
print(f"{step:>2}. result {m['latency_ms']:>6.0f} ms {m['result_preview'][:60]}")
cost = (prompt_tokens * PRICE_INPUT + completion_tokens * PRICE_OUTPUT) / 1_000_000
print(f" {prompt_tokens} in + {completion_tokens} out tokens = ${cost:.5f}")python explain_trace.py traces/shop-assistant-events.ndjson== q1
1. model 1832 ms 288 tokens -> check_stock
2. call check_stock {'sku': 'MS-30'}
3. result 951 ms {"sku": "MS-30", "units": 6}
4. model 937 ms 325 tokens -> final answer
570 in + 43 out tokens = $0.00023
== q2
1. model 1449 ms 352 tokens -> search_products
2. call search_products {'query': 'MS-10'}
3. result 89 ms [{"sku": "MS-10", "name": "MS-10 wireless mouse", "price_eur
4. model 1972 ms 485 tokens -> final answer
585 in + 252 out tokens = $0.00065
== q3
1. model 1181 ms 287 tokens -> search_products
2. call search_products {'query': 'ergonomic mouse'}
3. result 211 ms []
4. model 1551 ms 380 tokens -> search_products
5. call search_products {'query': 'mouse'}
6. result 231 ms []
7. model 2772 ms 630 tokens -> final answer
877 in + 420 out tokens = $0.00106Read q3 from the top and the problem is plain:
The model picked the right tool and passed a sensible query, ergonomic mouse.
The tool returned an empty list. This is the step that went wrong: the catalog has an "MS-30 ergonomic mouse".
The model tried again with a broader query, mouse, and got nothing again.
With two empty results, it told the customer the shop has no ergonomic mice.
The model did everything right with what it was given. Without the trace, you'd probably start rewriting the prompt. With it, you know the fix belongs in search_products, and you can see that the bug also cost an extra round trip: q3 used more than twice the tokens of q1.
The same reading works for most agent failures. Look at the first step where what happened differs from what you expected:
The model called the wrong tool, or none. Improve the tool's description.
The right tool got the wrong arguments. Describe the parameters better, or validate them in the server.
A tool returned wrong or empty data. Fix the tool, as here.
Everything was right, but the answer wasn't. Now it's the prompt or the model.
The same tool is called over and over. The agent is stuck, and you're paying for every loop.
Fix it and confirm with a new trace
Make the search match every word anywhere in the name:
words = query.lower().split()
# Match every word anywhere in the name, not just at the start.
return [
{"sku": sku, **p}
for sku, p in PRODUCTS.items()
if all(word in p["name"].lower() for word in words)
]Run shop_agent.py again in a clean folder, and q3 now gets it right:
[q3] Do you sell an ergonomic mouse?
Yes — we have the MS-30 ergonomic mouse (SKU: MS-30) for €49.00. There are 6 units in stock. Would you like to add one to your order or see more details?== q3
1. model 1482 ms 287 tokens -> search_products
2. call search_products {'query': 'ergonomic mouse'}
3. result 122 ms [{"sku": "MS-30", "name": "MS-30 ergonomic mouse", "price_eu
4. model 1384 ms 410 tokens -> check_stock
5. call check_stock {'sku': 'MS-30'}
6. result 1123 ms {"sku": "MS-30", "units": 6}
7. model 2410 ms 547 tokens -> final answer
948 in + 296 out tokens = $0.00083The same query now finds the product, and the agent goes on to check stock on its own. Keep the old and new trace side by side when you fix something: it's the quickest proof that the fix worked and didn't break the other questions.
[05]
Estimate the cost of each question
Promptise records tokens, not money. That's less of a gap than it sounds: prices change and differ between providers, so you multiply the tokens by your own rates anyway. explain_trace.py above does it with the OpenAI price list for gpt-5-mini at the time of writing, $0.25 per million input tokens and $2.00 per million output tokens.
Two things to know when you read the result:
Output tokens include reasoning. gpt-5-mini thinks before it answers, and OpenAI bills those reasoning tokens as output tokens. That's why the one-line answer to q2 used 252 output tokens.
Cached input is counted at full price. The trace doesn't record how many input tokens the provider served from its prompt cache, so this estimate is an upper bound.
For the whole session, get_stats() gives you the same totals: 2,032 input and 715 output tokens, about $0.0019 for three questions. Multiply by your daily traffic and you have a budget.
[06]
Send traces to your logging and tracing stack
A report on your laptop is fine for debugging. In production you want events where your team already looks. The same config can send them to several places at once:
import asyncio
import sys
from promptise import ObservabilityConfig, ObserveLevel, TransporterType, build_agent
from promptise.config import StdioServerSpec
async def main():
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
},
instructions="You are a shop assistant. Use your tools to answer. Be brief.",
observe=ObservabilityConfig(
level=ObserveLevel.STANDARD,
session_name="shop-assistant",
correlation_id="req-7f3a",
transporters=[
TransporterType.STRUCTURED_LOG,
TransporterType.WEBHOOK,
TransporterType.OTLP,
],
log_file="logs/agent.jsonl",
webhook_url="http://127.0.0.1:8231/events",
webhook_headers={"Authorization": "Bearer demo-token"},
otlp_endpoint="http://127.0.0.1:8230",
),
)
try:
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": "Is the MS-30 in stock?"}]}
)
print(result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())Structured logs go to logs/agent.jsonl, one JSON line per event, the format log pipelines such as ELK, Datadog or Splunk ingest. Tokens, latency, model and tool name are promoted to top-level fields, so you can filter and chart them without parsing. Here is the tool result from this run:
{"timestamp": "2026-10-10T16:47:33.535624+00:00", "level": "INFO", "service": "promptise", "session": "shop-assistant", "entry_id": "802122fa-e4bc-4a07-9f29-3de0f98796a4", "event_type": "tool.result", "category": "transparency", "message": "Tool completed: search_products", "agent_id": "agent", "duration_ms": 122.6, "correlation_id": "req-7f3a", "latency_ms": 122.6, "tool_name": "search_products", "metadata": {"result_preview": "[{\"sku\": \"MS-30\", \"name\": \"MS-30 ergonomic mouse\", \"price_eur\": 49.0}]", "run_id": "01a126b6-7220-7091-98b8-26ae4994a894"}}Webhooks receive each event as an HTTP POST with your headers. A small local endpoint on port 8231 printed what arrived:
Bearer… llm.start: LLM call started
Bearer… llm.end: LLM call completed (480 tokens, 2309.8ms)
Bearer… tool.call: Calling tool: search_products
Bearer… tool.result: Tool completed: search_products
…OpenTelemetry export uses OTLP over gRPC. session_name becomes the service name. For this guide, the spans went to a minimal OTLP/gRPC receiver on port 8230, built on the official opentelemetry-proto package, not to a full backend:
service=shop-assistant trace=ce7aea46 parent=- promptise.agent.llm.start span=0.03ms duration_ms=-
service=shop-assistant trace=740d9bb9 parent=- promptise.agent.llm.end span=0.06ms duration_ms=2309.830904006958
service=shop-assistant trace=1a99b495 parent=- promptise.agent.tool.call span=0.02ms duration_ms=-
service=shop-assistant trace=fef52e48 parent=- promptise.agent.tool.result span=0.05ms duration_ms=122.59101867675781
…Every event arrived, with its attributes under a promptise. prefix. Notice the shape, though. Each event is its own single-span trace with no parent, and each span lasts a fraction of a millisecond; the real duration is in the promptise.duration_ms attribute. Because no span has a parent, a tracing backend such as Jaeger can't group one request into a tree, and span durations won't mean much. A full backend wasn't tested for this guide. Treat the OTLP export as a stream of tagged events for now, not as nested traces.
Point otlp_endpoint at your own collector the same way, for example http://localhost:4317, the default. If the OpenTelemetry packages aren't installed, Promptise logs a warning and the OTLP transporter sends nothing, so check your logs after the first deploy.
[07]
Choose what gets recorded
Traces are only useful if you can keep them, so know what's in them. This is what 1.2.1 actually recorded for the same question at each setting, checked event by event:
Setting | What was recorded |
|---|---|
Default (STANDARD) | Model calls with tokens and latency, tool calls with arguments, tool results |
level=ObserveLevel.BASIC | The same, minus the llm.start events |
level=ObserveLevel.FULL | The same as STANDARD; it only adds a count of streamed tokens when you stream |
record_prompts=True | Adds the prompt and the model's response text, up to 2,000 characters each |
level=ObserveLevel.OFF | The same as STANDARD; it does not turn recording off |
So the levels change less than the docs suggest. Prompt and response text depend on record_prompts, not on the level. To turn observability off, leave out observe.
[08]
Honest limits
Promptise's agent observability records the right things. A few parts around it aren't finished in 1.2.1, and you'll want to know before you build on them:
The HTML report's stat cards and filters don't work. They show zeros, as described in step 3. The embedded data is correct.
`generate_report(path)` changes the file name. It writes <name>-report-<timestamp>.html in the folder of path, returns path anyway, and ignores its title argument.
Tool errors don't count as errors. An MCP server reports a failed tool call as a result, so it's recorded as tool.result and error_count stays at 0. Search result_preview for "error" to find these. There's an example below.
There's no event for the start and end of a request. The docs list agent input and output events, but none were recorded in these runs. Tag requests with CallerContext as shown above to tell them apart.
Don't combine `observe` with the semantic cache yet. With both on, cache hits are thrown away and every question goes to the model again; the agent logs
Cache check failed, continuing without cache. If you need the cache, leave observability off for that agent until this is fixed.OpenTelemetry spans are flat. One trace per event, no parent-child links, and durations in an attribute instead of the span itself. The spans also don't carry user_id, session_id or correlation_id.
The in-memory timeline is capped. It keeps the newest 100,000 events by default (max_entries), and older ones are dropped. Long-running agents should stream to the JSON, log or webhook transporters.
Here is the failed tool call from the first point, a lookup of an SKU that doesn't exist:
tool.call {'sku': 'XX-999'}
tool.result {
"error": {
"code": "TOOL_ERROR",
"message": "Unknown SKU XX-999.",
…
error_count: 0[09]
Frequently asked questions
What is AI agent observability?
It's the record of everything an agent did to produce an answer: each model call with its tokens and latency, each tool call with its arguments and result, and any errors. Logs tell you that something happened; a trace of the steps tells you why the agent decided what it decided, which is what you need to debug it or explain it to someone else.
Does Promptise track the cost of LLM calls?
It tracks tokens per call, split into input and output, plus the exact model version. It doesn't convert them to money. Multiply by your provider's prices, as explain_trace.py does, so the numbers stay right when prices change.
Is there an open source alternative to LangSmith for tracing agents?
If you build with Promptise Foundry, its observability is part of the open source framework (Apache 2.0) and needs no account or hosted service: traces stay in local files unless you send them somewhere. It doesn't give you a hosted trace UI like LangSmith. For a shared view, send events to tools you already run, through structured logs, webhooks or OpenTelemetry.
Will customer data end up in my traces?
Tool arguments and tool results will, up to 2,000 characters each, even with the default settings. Prompts and model responses are only recorded with record_prompts=True. Plan retention and access for trace files and log streams as you would for any store of customer data.
Can I send AI agent traces to OpenTelemetry?
Yes. Add TransporterType.OTLP, set otlp_endpoint, and install the OpenTelemetry SDK and OTLP gRPC exporter. In 1.2.1 each event arrives as its own span with promptise.* attributes, so it works well for search and dashboards but doesn't yet show a request as one nested trace.
[10]
Where to go next
Observability: every option of ObservabilityConfig and the transporters.
Observability API reference: ObservabilityCollector, TimelineEntry and the transporter classes.
Building Agents: build_agent, CallerContext and the rest of the agent API.
Events & Notifications: alerts for slow tools and errors, separate from the trace.
Observability & Monitoring for MCP servers: the other side of every tool call.
How to Connect MCP Servers to Your AI Agent in Python: the agent setup this guide builds on.
Build a Production MCP Server in Python: audit logs and server-side limits for the tools your agent calls.