PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

AI Agent Observability: Trace Tool Calls, Tokens and Cost

Trace every LLM call and tool call your AI agent makes, with tokens, latency and cost, then use the trace to find the exact step that went wrong.

Level
Intermediate
Reading time
18 min
Published
Oct 10, 2026
By
Promptise Team
  • AI Agents
  • Observability
  • OpenTelemetry
  • Tool Calling
  • Debugging
  • Promptise Foundry

An agent gives a customer a wrong answer, and you have to explain why. Was it the prompt, the model, or one of its tools? Without a trace you're guessing. AI agent observability means recording every LLM call and every tool call, with its arguments, result, tokens and latency, so you can replay what the agent did and point at the step that went wrong. In this guide you'll add it to a small shop assistant with two MCP tools, read the real trace, find a real bug with it, estimate cost per question, and send the same events to a log file, a webhook and OpenTelemetry. Every snippet ran against Promptise Foundry 1.2.1, and the output is what it printed.

[01]

How do you add observability to an AI agent?

With Promptise Foundry, pass observe=True to build_agent. From then on every model call (tokens, latency, model version, which tools it asked for) and every tool call (arguments, result, latency) is recorded on a timeline. agent.get_stats() returns the totals, and an HTML report is written when the agent shuts down:

Python
1
2
3
4
5
6
7
8
9
10
11
12
13
14
from promptise import build_agent
from promptise.config import StdioServerSpec

agent = await build_agent(
    model="openai:gpt-5-mini",
    servers={
        "shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
    },
    instructions="You are a shop assistant. Use your tools to answer. Be brief.",
    observe=True,
)
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
print(json.dumps(agent.get_stats(), indent=2))
print(agent.generate_report("shop-report.html"))

No trace data leaves your machine unless you add a transporter that sends it. The trace lives in memory and in local files, and the code that records it is open source. The rest of this guide shows what the trace contains and how to use it.


[02]

How agent observability works

When observability is on, build_agent attaches a PromptiseCallbackHandler to every run. It turns each model call and tool call into a timeline event and hands it to an ObservabilityCollector. The collector keeps the events in memory and passes each one to the transporters you chose.

Rendering diagram…

These are the events you'll see from an agent with MCP tools, and what each one carries:

Event

Recorded when

What it carries

llm.start

A model call begins

Model name

llm.end

A model call returns

Prompt, completion and total tokens, latency, exact model version, the tools the model asked for

tool.call

A tool starts

Tool name and arguments

tool.result

A tool returns

The first 2,000 characters of the result, latency

llm.error

A model call raises

Error type, message and traceback

Each event can also carry a user_id and session_id, taken from the CallerContext you pass to ainvoke. That turns out to be the easiest way to pull one request back out of a busy trace.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry: pip install promptise. Everything up to the OpenTelemetry section needs nothing else.

  • For OpenTelemetry export: pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc. The docs suggest pip install "promptise[all]", which works too but also installs much more, including PyTorch.

  • An API key for a model provider. This guide uses OpenAI, so set OPENAI_API_KEY.


[04]

Trace an AI agent, step by step

You'll build a shop assistant that answers questions about products and stock, run three real questions through it, and use the trace to find out why one answer is wrong.

Step 01

Give the agent two tools

The tools live in a small MCP server. check_stock waits 0.8 seconds on purpose, standing in for a slow warehouse API, so you can see latency in the trace. search_products has a bug, the kind that slips into real code: it only matches the start of the product name, like a LIKE 'query%' database index.

Pythonshop_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
import asyncio

from promptise.mcp.server import MCPServer, ToolError

server = MCPServer("shop")

# A stand-in for your product catalog and warehouse system.
PRODUCTS = {
    "KB-200": {"name": "KB-200 mechanical keyboard", "price_eur": 89.0},
    "MS-10": {"name": "MS-10 wireless mouse", "price_eur": 24.0},
    "MS-30": {"name": "MS-30 ergonomic mouse", "price_eur": 49.0},
}
STOCK = {"KB-200": 14, "MS-10": 0, "MS-30": 6}


@server.tool()
async def search_products(query: str) -> list[dict]:
    """Search the catalog by product name. Returns SKU, name and price in EUR.

    Args:
        query: Words to look for in the product name, for example "mouse".
    """
    q = query.lower()
    # Prefix match, like a LIKE 'query%' index in a database.
    return [{"sku": sku, **p} for sku, p in PRODUCTS.items() if p["name"].lower().startswith(q)]


@server.tool()
async def check_stock(sku: str) -> dict:
    """Check how many units of a product are in the warehouse.

    Args:
        sku: The product SKU, for example "KB-200".
    """
    await asyncio.sleep(0.8)  # the warehouse API is slow
    if sku.upper() not in STOCK:
        raise ToolError(f"Unknown SKU {sku}.")
    return {"sku": sku.upper(), "units": STOCK[sku.upper()]}


if __name__ == "__main__":
    server.run()

If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python walks through this part.

Step 02

Turn on observability and run real questions

observe=True is enough to start. For a trace you can come back to, pass an ObservabilityConfig instead. This one keeps the HTML report and adds the JSON transporter, which streams every event to a file as it happens. The CallerContext tags each question's events with its own session_id.

Pythonshop_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
import asyncio
import json
import sys

from promptise import CallerContext, ObservabilityConfig, TransporterType, build_agent
from promptise.config import StdioServerSpec

QUESTIONS = {
    "q1": "Is the MS-30 in stock?",
    "q2": "How much is the MS-10?",
    "q3": "Do you sell an ergonomic mouse?",
}


async def main():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
        },
        instructions="You are a shop assistant. Use your tools to answer. Be brief.",
        observe=ObservabilityConfig(
            session_name="shop-assistant",
            transporters=[TransporterType.HTML, TransporterType.JSON],
            output_dir="traces",
        ),
    )
    try:
        for qid, question in QUESTIONS.items():
            # Tag every event of this request, so you can pull it out of the trace later.
            caller = CallerContext(user_id="customer-42", metadata={"session_id": qid})
            result = await agent.ainvoke(
                {"messages": [{"role": "user", "content": question}]}, caller=caller
            )
            print(f"[{qid}] {question}")
            print(f"      {result['messages'][-1].content}")

        print(json.dumps(agent.get_stats(), indent=2))
    finally:
        await agent.shutdown()


asyncio.run(main())
>_Terminal
python shop_agent.py
Output
[q1] Is the MS-30 in stock?
      Yes — we have 6 units of the MS-30 in stock.
[q2] How much is the MS-10?
      The MS-10 wireless mouse costs €24.00. Would you like me to check stock or add one to your cart?
[q3] Do you sell an ergonomic mouse?
      I don’t see any ergonomic mice in our catalog right now. Would you like me to search for related items (vertical mice, trackballs, or other ergonomic pointing devices) or check for a specific brand/model?
{
  "entry_count": 22,
  "agent_count": 1,
  "total_duration_s": 13.86664605140686,
  "total_prompt_tokens": 2032,
  "total_completion_tokens": 715,
  "total_tokens": 2747,
  "llm_call_count": 7,
  "tool_call_count": 4,
  "error_count": 0,
  "retry_count": 0,
  "latency_p50_ms": 1180.9399127960205,
  "latency_p95_ms": 2771.5067863464355,
  "latency_p99_ms": 2771.5067863464355,
…
  "events_by_type": {
    "llm.start": 7,
    "llm.end": 7,
    "tool.call": 4,
    "tool.result": 4
  }
}

Three questions took 7 model calls, 4 tool calls and 2,747 tokens. Two details help you read these numbers. The latency percentiles mix model calls and tool calls, because they cover every event that has a duration. And total_duration_s counts from when the agent was built, not just the time spent answering.

The third answer is wrong. The catalog has an ergonomic mouse, the MS-30, and the agent said it didn't. The trace will tell you why.

Step 03

Open the HTML report

When the agent shut down, the HTML transporter wrote traces/shop-assistant-report-20261010_184106.html. It's one self-contained file, with no external scripts or styles, that you can open in any browser or attach to a ticket. Rendered in a browser (here, headless Chrome), it shows six stat cards, filter buttons and one row per event:

Output
title: Promptise Agent Report
…
rendered stat cards: [('22', 'Total Events'), ('0', 'Total Tokens'), ('0', 'LLM Calls'), ('0', 'Tool Calls'), ('0', 'Errors'), ('0', 'Cache Hits')]
rendered filter buttons: ['All', 'Tool', 'Llm', 'Error', 'Cache']
rendered rows: 22
   📌 llm.start LLM call started 18:40:50
   📌 llm.end LLM call completed (288 tokens, 1831.7ms) 18:40:52
   📌 tool.call Calling tool: check_stock 18:40:52
   📌 tool.result Tool completed: check_stock 18:40:53

Each row shows the event type, a one-line description and the time. The token count and latency of each model call are in its description.

Warning

In 1.2.1 the report's stat cards show 0 tokens, 0 LLM calls and 0 tool calls, and the Tool and LLM filters hide every row. The page looks for event names like llm_end, while the events are named llm.end. The data itself is right: the file embeds the full trace as JSON, including "total_tokens": 2747, every tool's arguments and every result. Use get_stats() or the JSON files for numbers until this is fixed.

You can also write a report at any point with agent.generate_report(path). Be aware that it doesn't use your path as the file name. From the quick example at the top:

Python
print(agent.generate_report("shop-report.html"))
Output
shop-report.html
>_Terminal
ls *.html reports/
Output
shop-report-report-20261010_185045.html

reports/:
promptise-report-20261010_185048.html

It printed shop-report.html, but the file it wrote has a timestamp added. The second file is the automatic report from observe=True, which goes to ./reports unless you set output_dir.

Step 04

Find the step where the agent went wrong

The JSON transporter streamed every event to traces/shop-assistant-events.ndjson, one JSON object per line. Each line has the event type, the session_id you set, and a metadata object with the details. This short script turns that file into numbered steps per question, with the token cost of each:

Pythonexplain_trace.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
"""Print a saved Promptise trace as numbered steps, one block per request."""

import json
import sys
from collections import defaultdict

# USD per million tokens for openai:gpt-5-mini. Check your provider's current prices.
PRICE_INPUT = 0.25
PRICE_OUTPUT = 2.00

requests = defaultdict(list)
with open(sys.argv[1]) as f:
    for line in f:
        event = json.loads(line)
        requests[event["session_id"]].append(event)

for request_id, events in requests.items():
    print(f"== {request_id}")
    step, prompt_tokens, completion_tokens = 0, 0, 0
    for e in events:
        m = e["metadata"]
        if e["event_type"] == "llm.end":
            step += 1
            prompt_tokens += m["prompt_tokens"]
            completion_tokens += m["completion_tokens"]
            wants = ", ".join(m.get("tool_calls", [])) or "final answer"
            print(f"{step:>2}. model   {m['latency_ms']:>6.0f} ms  {m['total_tokens']:>4} tokens -> {wants}")
        elif e["event_type"] == "tool.call":
            step += 1
            print(f"{step:>2}. call    {m['tool_name']} {m['arguments']}")
        elif e["event_type"] == "tool.result":
            step += 1
            print(f"{step:>2}. result  {m['latency_ms']:>6.0f} ms  {m['result_preview'][:60]}")
    cost = (prompt_tokens * PRICE_INPUT + completion_tokens * PRICE_OUTPUT) / 1_000_000
    print(f"    {prompt_tokens} in + {completion_tokens} out tokens = ${cost:.5f}")
>_Terminal
python explain_trace.py traces/shop-assistant-events.ndjson
Output
== q1
 1. model     1832 ms   288 tokens -> check_stock
 2. call    check_stock {'sku': 'MS-30'}
 3. result     951 ms  {"sku": "MS-30", "units": 6}
 4. model      937 ms   325 tokens -> final answer
    570 in + 43 out tokens = $0.00023
== q2
 1. model     1449 ms   352 tokens -> search_products
 2. call    search_products {'query': 'MS-10'}
 3. result      89 ms  [{"sku": "MS-10", "name": "MS-10 wireless mouse", "price_eur
 4. model     1972 ms   485 tokens -> final answer
    585 in + 252 out tokens = $0.00065
== q3
 1. model     1181 ms   287 tokens -> search_products
 2. call    search_products {'query': 'ergonomic mouse'}
 3. result     211 ms  []
 4. model     1551 ms   380 tokens -> search_products
 5. call    search_products {'query': 'mouse'}
 6. result     231 ms  []
 7. model     2772 ms   630 tokens -> final answer
    877 in + 420 out tokens = $0.00106

Read q3 from the top and the problem is plain:

  1. The model picked the right tool and passed a sensible query, ergonomic mouse.

  2. The tool returned an empty list. This is the step that went wrong: the catalog has an "MS-30 ergonomic mouse".

  3. The model tried again with a broader query, mouse, and got nothing again.

  4. With two empty results, it told the customer the shop has no ergonomic mice.

The model did everything right with what it was given. Without the trace, you'd probably start rewriting the prompt. With it, you know the fix belongs in search_products, and you can see that the bug also cost an extra round trip: q3 used more than twice the tokens of q1.

The same reading works for most agent failures. Look at the first step where what happened differs from what you expected:

  • The model called the wrong tool, or none. Improve the tool's description.

  • The right tool got the wrong arguments. Describe the parameters better, or validate them in the server.

  • A tool returned wrong or empty data. Fix the tool, as here.

  • Everything was right, but the answer wasn't. Now it's the prompt or the model.

  • The same tool is called over and over. The agent is stuck, and you're paying for every loop.

Step 05

Fix it and confirm with a new trace

Make the search match every word anywhere in the name:

Pythonshop_server.py
1
2
3
4
5
6
7
    words = query.lower().split()
    # Match every word anywhere in the name, not just at the start.
    return [
        {"sku": sku, **p}
        for sku, p in PRODUCTS.items()
        if all(word in p["name"].lower() for word in words)
    ]

Run shop_agent.py again in a clean folder, and q3 now gets it right:

Output
[q3] Do you sell an ergonomic mouse?
      Yes — we have the MS-30 ergonomic mouse (SKU: MS-30) for €49.00. There are 6 units in stock. Would you like to add one to your order or see more details?
Output
== q3
 1. model     1482 ms   287 tokens -> search_products
 2. call    search_products {'query': 'ergonomic mouse'}
 3. result     122 ms  [{"sku": "MS-30", "name": "MS-30 ergonomic mouse", "price_eu
 4. model     1384 ms   410 tokens -> check_stock
 5. call    check_stock {'sku': 'MS-30'}
 6. result    1123 ms  {"sku": "MS-30", "units": 6}
 7. model     2410 ms   547 tokens -> final answer
    948 in + 296 out tokens = $0.00083

The same query now finds the product, and the agent goes on to check stock on its own. Keep the old and new trace side by side when you fix something: it's the quickest proof that the fix worked and didn't break the other questions.

Note

The JSON transporter appends to shop-assistant-events.ndjson on every run. Delete or rotate it between runs, or give each run its own session_name, so old events don't mix into new ones.


[05]

Estimate the cost of each question

Promptise records tokens, not money. That's less of a gap than it sounds: prices change and differ between providers, so you multiply the tokens by your own rates anyway. explain_trace.py above does it with the OpenAI price list for gpt-5-mini at the time of writing, $0.25 per million input tokens and $2.00 per million output tokens.

Two things to know when you read the result:

  • Output tokens include reasoning. gpt-5-mini thinks before it answers, and OpenAI bills those reasoning tokens as output tokens. That's why the one-line answer to q2 used 252 output tokens.

  • Cached input is counted at full price. The trace doesn't record how many input tokens the provider served from its prompt cache, so this estimate is an upper bound.

For the whole session, get_stats() gives you the same totals: 2,032 input and 715 output tokens, about $0.0019 for three questions. Multiply by your daily traffic and you have a budget.


[06]

Send traces to your logging and tracing stack

A report on your laptop is fine for debugging. In production you want events where your team already looks. The same config can send them to several places at once:

Pythonexport_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
import asyncio
import sys

from promptise import ObservabilityConfig, ObserveLevel, TransporterType, build_agent
from promptise.config import StdioServerSpec


async def main():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "shop": StdioServerSpec(command=sys.executable, args=["shop_server.py"]),
        },
        instructions="You are a shop assistant. Use your tools to answer. Be brief.",
        observe=ObservabilityConfig(
            level=ObserveLevel.STANDARD,
            session_name="shop-assistant",
            correlation_id="req-7f3a",
            transporters=[
                TransporterType.STRUCTURED_LOG,
                TransporterType.WEBHOOK,
                TransporterType.OTLP,
            ],
            log_file="logs/agent.jsonl",
            webhook_url="http://127.0.0.1:8231/events",
            webhook_headers={"Authorization": "Bearer demo-token"},
            otlp_endpoint="http://127.0.0.1:8230",
        ),
    )
    try:
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Is the MS-30 in stock?"}]}
        )
        print(result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())

Structured logs go to logs/agent.jsonl, one JSON line per event, the format log pipelines such as ELK, Datadog or Splunk ingest. Tokens, latency, model and tool name are promoted to top-level fields, so you can filter and chart them without parsing. Here is the tool result from this run:

Output
{"timestamp": "2026-10-10T16:47:33.535624+00:00", "level": "INFO", "service": "promptise", "session": "shop-assistant", "entry_id": "802122fa-e4bc-4a07-9f29-3de0f98796a4", "event_type": "tool.result", "category": "transparency", "message": "Tool completed: search_products", "agent_id": "agent", "duration_ms": 122.6, "correlation_id": "req-7f3a", "latency_ms": 122.6, "tool_name": "search_products", "metadata": {"result_preview": "[{\"sku\": \"MS-30\", \"name\": \"MS-30 ergonomic mouse\", \"price_eur\": 49.0}]", "run_id": "01a126b6-7220-7091-98b8-26ae4994a894"}}

Webhooks receive each event as an HTTP POST with your headers. A small local endpoint on port 8231 printed what arrived:

Output
Bearer… llm.start: LLM call started
Bearer… llm.end: LLM call completed (480 tokens, 2309.8ms)
Bearer… tool.call: Calling tool: search_products
Bearer… tool.result: Tool completed: search_products
…

OpenTelemetry export uses OTLP over gRPC. session_name becomes the service name. For this guide, the spans went to a minimal OTLP/gRPC receiver on port 8230, built on the official opentelemetry-proto package, not to a full backend:

Output
service=shop-assistant trace=ce7aea46 parent=- promptise.agent.llm.start span=0.03ms duration_ms=-
service=shop-assistant trace=740d9bb9 parent=- promptise.agent.llm.end span=0.06ms duration_ms=2309.830904006958
service=shop-assistant trace=1a99b495 parent=- promptise.agent.tool.call span=0.02ms duration_ms=-
service=shop-assistant trace=fef52e48 parent=- promptise.agent.tool.result span=0.05ms duration_ms=122.59101867675781
…

Every event arrived, with its attributes under a promptise. prefix. Notice the shape, though. Each event is its own single-span trace with no parent, and each span lasts a fraction of a millisecond; the real duration is in the promptise.duration_ms attribute. Because no span has a parent, a tracing backend such as Jaeger can't group one request into a tree, and span durations won't mean much. A full backend wasn't tested for this guide. Treat the OTLP export as a stream of tagged events for now, not as nested traces.

Point otlp_endpoint at your own collector the same way, for example http://localhost:4317, the default. If the OpenTelemetry packages aren't installed, Promptise logs a warning and the OTLP transporter sends nothing, so check your logs after the first deploy.

Tip

correlation_id appears only in the structured log lines. To find one request in webhook payloads or the JSON files, use the user_id and session_id from CallerContext, which every event there carries. The OpenTelemetry spans carry neither in 1.2.1.


[07]

Choose what gets recorded

Traces are only useful if you can keep them, so know what's in them. This is what 1.2.1 actually recorded for the same question at each setting, checked event by event:

Setting

What was recorded

Default (STANDARD)

Model calls with tokens and latency, tool calls with arguments, tool results

level=ObserveLevel.BASIC

The same, minus the llm.start events

level=ObserveLevel.FULL

The same as STANDARD; it only adds a count of streamed tokens when you stream

record_prompts=True

Adds the prompt and the model's response text, up to 2,000 characters each

level=ObserveLevel.OFF

The same as STANDARD; it does not turn recording off

So the levels change less than the docs suggest. Prompt and response text depend on record_prompts, not on the level. To turn observability off, leave out observe.

Important

record_prompts=False keeps prompts out of the trace, but tool arguments and tool results are always recorded, up to 2,000 characters each. If your tools return personal data, it will be in your trace files, logs and webhook payloads. Keep sensitive fields out of tool results, or handle trace storage like the data it contains. For per-user cleanup, the collector has purge_user(user_id), which removes that user's events from memory but not from files already written.


[08]

Honest limits

Promptise's agent observability records the right things. A few parts around it aren't finished in 1.2.1, and you'll want to know before you build on them:

  • The HTML report's stat cards and filters don't work. They show zeros, as described in step 3. The embedded data is correct.

  • `generate_report(path)` changes the file name. It writes <name>-report-<timestamp>.html in the folder of path, returns path anyway, and ignores its title argument.

  • Tool errors don't count as errors. An MCP server reports a failed tool call as a result, so it's recorded as tool.result and error_count stays at 0. Search result_preview for "error" to find these. There's an example below.

  • There's no event for the start and end of a request. The docs list agent input and output events, but none were recorded in these runs. Tag requests with CallerContext as shown above to tell them apart.

  • Don't combine `observe` with the semantic cache yet. With both on, cache hits are thrown away and every question goes to the model again; the agent logs Cache check failed, continuing without cache. If you need the cache, leave observability off for that agent until this is fixed.

  • OpenTelemetry spans are flat. One trace per event, no parent-child links, and durations in an attribute instead of the span itself. The spans also don't carry user_id, session_id or correlation_id.

  • The in-memory timeline is capped. It keeps the newest 100,000 events by default (max_entries), and older ones are dropped. Long-running agents should stream to the JSON, log or webhook transporters.

Here is the failed tool call from the first point, a lookup of an SKU that doesn't exist:

Output
tool.call {'sku': 'XX-999'}
tool.result {
  "error": {
    "code": "TOOL_ERROR",
    "message": "Unknown SKU XX-999.",
…
error_count: 0

[09]

Frequently asked questions

What is AI agent observability?

It's the record of everything an agent did to produce an answer: each model call with its tokens and latency, each tool call with its arguments and result, and any errors. Logs tell you that something happened; a trace of the steps tells you why the agent decided what it decided, which is what you need to debug it or explain it to someone else.

Does Promptise track the cost of LLM calls?

It tracks tokens per call, split into input and output, plus the exact model version. It doesn't convert them to money. Multiply by your provider's prices, as explain_trace.py does, so the numbers stay right when prices change.

Is there an open source alternative to LangSmith for tracing agents?

If you build with Promptise Foundry, its observability is part of the open source framework (Apache 2.0) and needs no account or hosted service: traces stay in local files unless you send them somewhere. It doesn't give you a hosted trace UI like LangSmith. For a shared view, send events to tools you already run, through structured logs, webhooks or OpenTelemetry.

Will customer data end up in my traces?

Tool arguments and tool results will, up to 2,000 characters each, even with the default settings. Prompts and model responses are only recorded with record_prompts=True. Plan retention and access for trace files and log streams as you would for any store of customer data.

Can I send AI agent traces to OpenTelemetry?

Yes. Add TransporterType.OTLP, set otlp_endpoint, and install the OpenTelemetry SDK and OTLP gRPC exporter. In 1.2.1 each event arrives as its own span with promptise.* attributes, so it works well for search and dashboards but doesn't yet show a request as one nested trace.


[10]

Where to go next

  • Observability: every option of ObservabilityConfig and the transporters.

  • Observability API reference: ObservabilityCollector, TimelineEntry and the transporter classes.

  • Building Agents: build_agent, CallerContext and the rest of the agent API.

  • Events & Notifications: alerts for slow tools and errors, separate from the trace.

  • Observability & Monitoring for MCP servers: the other side of every tool call.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent setup this guide builds on.

  • Build a Production MCP Server in Python: audit logs and server-side limits for the tools your agent calls.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths