PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

MCP Server Monitoring: Prometheus, OpenTelemetry and Logs

Monitor an MCP server in Python: Prometheus metrics, JSON logs, OpenTelemetry traces and alerts that show which tools are slow, failing or abused.

Level
Intermediate
Reading time
20 min
Published
Oct 10, 2026
By
Promptise Team
  • MCP
  • MCP Server
  • Observability
  • Prometheus
  • OpenTelemetry
  • Promptise Foundry

Once real agents call your MCP server, three questions come up fast: which tool is slow, which one is failing, and who is calling it far more than they should. MCP server monitoring answers them with three signals: metrics for trends and alerts, structured logs for the details of each call, and traces for where the time went. In this guide you'll build a small inventory server in Python with one slow tool and one flaky one, send it real traffic from a well-behaved agent and a misbehaving bot, and read all three signals, plus a live terminal dashboard. Every snippet was run against Promptise Foundry 1.2.1, and every output is what it printed.

[01]

How do you monitor an MCP server?

Wrap every tool call in middleware that records it, then send the records to the tools your team already uses. With Promptise Foundry, two lines of middleware give you JSON logs and Prometheus metrics, and one call from prometheus_client serves them for scraping:

Pythonquickstart_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import logging

from prometheus_client import start_http_server

from promptise.mcp.server import MCPServer, PrometheusMiddleware, StructuredLoggingMiddleware

logging.basicConfig(level=logging.INFO, format="%(message)s")

server = MCPServer("inventory")
server.add_middleware(StructuredLoggingMiddleware())  # one JSON log line per call start and end
server.add_middleware(PrometheusMiddleware())  # call counts, errors, latency histogram


@server.tool()
async def check_stock(sku: str) -> dict:
    """How many units of a product are in the warehouse right now."""
    return {"sku": sku, "units": 42}


if __name__ == "__main__":
    start_http_server(8411, addr="127.0.0.1")  # Prometheus scrapes http://127.0.0.1:8411/metrics
    server.run(transport="http", host="127.0.0.1", port=8410)

After one tool call, the server logged two JSON lines and Prometheus could scrape the call:

Output
{"event": "tool_call_start", "tool": "check_stock", "request_id": "e0260b6f7d1b"}
{"event": "tool_call_end", "tool": "check_stock", "request_id": "e0260b6f7d1b", "duration_ms": 0.25, "status": "ok"}
Output
$ curl -s http://127.0.0.1:8411/metrics | grep "^mcp_"
mcp_tool_calls_total{status="ok",tool="check_stock"} 1.0
…
mcp_tool_duration_seconds_count{tool="check_stock"} 1.0
mcp_tool_duration_seconds_sum{tool="check_stock"} 5.179100116947666e-05
…
mcp_tool_in_flight{tool="check_stock"} 0.0

The metrics live on their own port because the MCP server's HTTP port only serves /mcp. More on that in step 2. The rest of this guide adds tracing, health checks, per-client numbers and alerts, and puts them under real load.


[02]

How it works

Every tool call passes through the server's middleware in the order you add it. Each observability middleware times the call, notes whether it raised, and hands the result to its own destination:

Rendering diagram…

Authentication goes first so every later middleware knows the caller. Each piece answers a different question:

Question

Signal

Promptise piece

Which tool is slow?

Latency histogram, spans

PrometheusMiddleware, OTelMiddleware

Which tool is failing, and why?

Error counter, log lines with the error text

PrometheusMiddleware, StructuredLoggingMiddleware

Who is hammering the server?

client_id on logs and spans, a per-client counter

StructuredLoggingMiddleware, OTelMiddleware, six lines of your own

Is it up and ready?

Liveness and readiness probes

HealthCheck

What's happening right now?

A live terminal UI

server.run(dashboard=True)

Who did what, provably?

A signed audit log

AuditMiddleware


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry plus the Prometheus and OpenTelemetry packages. They aren't installed with plain pip install promptise, and without them PrometheusMiddleware() and OTelMiddleware() raise ImportError when you create them. pip install "promptise[all]" also pulls them in, along with everything else.

  • No model API key. The traffic in this guide comes from a small script, so you can rerun it as often as you like.

  • Optionally jq for querying the logs, and a Prometheus server and an OpenTelemetry Collector if you want to see the data in your own stack.

>_Terminal
pip install promptise prometheus-client opentelemetry-sdk opentelemetry-exporter-otlp-proto-http

This guide ran with prometheus-client 0.26.0 and OpenTelemetry 1.45.1.


[04]

Monitor the inventory MCP server, step by step

The server has three tools. check_stock is fast. forecast_demand is slow, between 0.8 and 2.5 seconds. reserve_stock fails about one call in four, like a flaky warehouse API. Two API keys identify the callers: shop-agent behaves, report-bot doesn't. The complete server file is at the end of this section; each step shows the part it adds.

Step 01

Build the server and know who is calling

Monitoring per caller only works if the server knows the caller. AuthMiddleware with APIKeyAuth turns each key into a client_id, and because it's added first, every middleware after it can read that ID:

Pythoninventory_server.py
server = MCPServer("inventory", version="1.0.0", require_auth=True)
Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
server.add_middleware(
    AuthMiddleware(
        APIKeyAuth(
            keys={
                os.environ["SHOP_AGENT_KEY"]: {"client_id": "shop-agent", "roles": ["agent"]},
                os.environ["REPORT_BOT_KEY"]: {"client_id": "report-bot", "roles": ["agent"]},
            }
        )
    )
)

The tools fail the way real tools do, with a ToolError the agent can read:

Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
@server.tool(read_only_hint=True)
async def forecast_demand(sku: str, weeks: int = 4) -> dict:
    """Forecast how many units of a product will sell in the next few weeks."""
    with tracer.start_as_current_span("load_sales_history") as span:
        span.set_attribute("inventory.sku", sku)
        await asyncio.sleep(random.uniform(0.8, 2.5))  # stands in for a slow query
    return {"sku": sku, "weeks": weeks, "expected_units": 6 * weeks}


@server.tool()
async def reserve_stock(sku: str, units: int) -> dict:
    """Hold units of a product for an order."""
    if random.random() < 0.25:  # stands in for a flaky warehouse API
        raise ToolError("The warehouse API did not answer. Try again in a moment.", retryable=True)
    if STOCK.get(sku, 0) < units:
        raise ToolError(f"Only {STOCK.get(sku, 0)} units of {sku} are left.")
    STOCK[sku] -= units
    return {"sku": sku, "reserved": units, "left": STOCK[sku]}

The span inside forecast_demand comes into play in step 4.

Step 02

Expose Prometheus metrics

PrometheusMiddleware records three metrics for every call: a counter mcp_tool_calls_total with tool and status labels, a histogram mcp_tool_duration_seconds per tool, and a gauge mcp_tool_in_flight per tool. Pass namespace="myapp" to change the mcp_ prefix.

Prometheus metrics answer "which tool", but not "which caller". A six-line middleware of your own adds a counter with a client label. It sits before the rate limiter, so calls the limiter rejects are counted too, which is exactly what you want when someone floods you:

Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
11
CALLS_BY_CLIENT = Counter(
    "mcp_tool_calls_by_client_total",
    "MCP tool calls per authenticated client, including rejected ones",
    ["client", "tool"],
)


async def count_calls_per_client(ctx, call_next):
    """Prometheus counter with a client label, so you can see who is calling."""
    CALLS_BY_CLIENT.labels(client=ctx.client_id or "anonymous", tool=ctx.tool_name).inc()
    return await call_next(ctx)

Then serve the metrics. The docstring of PrometheusMiddleware says they're available at GET /metrics on the HTTP transport, but in 1.2.1 that route doesn't exist. With a valid key, the MCP port answers:

Output
$ curl -i -H "x-api-key: $SHOP_AGENT_KEY" http://127.0.0.1:8410/metrics
HTTP/1.1 404 Not Found
…
Not Found

So start prometheus_client's own exporter on a second port before the server runs. It serves the default registry from a background thread:

Pythoninventory_server.py
if __name__ == "__main__":
    start_http_server(8411, addr="127.0.0.1")  # Prometheus scrapes http://127.0.0.1:8411/metrics
    server.run(transport="http", host="127.0.0.1", port=8410, dashboard="--dashboard" in sys.argv)

A separate port has an upside: the MCP port requires an API key, and Prometheus doesn't need one. Keep the metrics port on your internal network.

Note

The Promptise docs also show returning prom.get_metrics_text() from an MCP resource. That's readable by agents, but Prometheus can't scrape an MCP resource, so use a port for your monitoring stack.

Step 03

Log every call as JSON

StructuredLoggingMiddleware writes two entries per call to the promptise.server logger: tool_call_start, and tool_call_end with duration_ms, status, the error message on failure, and the client_id. Each entry is a JSON string, but two things stop it from being a clean JSON log on its own: the entries carry no timestamp, and Promptise writes plain-text messages to the same logger, such as MCP server running on http://127.0.0.1:8410/mcp. A small formatter fixes both, so every line your log shipper reads is one JSON object:

Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
class JsonLines(logging.Formatter):
    """Every record becomes one JSON object. Tool-call entries are JSON already."""

    def format(self, record):
        message = record.getMessage()
        try:
            entry = json.loads(message)
        except ValueError:
            entry = {"event": "log", "message": message}
        when = datetime.fromtimestamp(record.created, tz=timezone.utc).isoformat(timespec="milliseconds")
        entry = {"time": when, "level": record.levelname, **entry}
        if record.exc_info:
            entry["exception"] = self.formatException(record.exc_info)
        return json.dumps(entry)


json_lines = logging.StreamHandler(sys.stderr)
json_lines.setFormatter(JsonLines())
server_log = logging.getLogger("promptise.server")
server_log.addHandler(json_lines)
server_log.setLevel(logging.INFO)
server_log.propagate = False

Uvicorn keeps writing its own plain-text access log to stdout, which is useful in its own right, as you'll see under the honest limits.

Warning

StructuredLoggingMiddleware(include_args=True) is accepted but has no effect in 1.2.1: the log lines never contain the arguments. If you need arguments on record, use AuditMiddleware(include_args=True), which does write them.

Step 04

Trace every call with OpenTelemetry

Metrics tell you a tool is slow. A trace tells you which part of it is slow. OTelMiddleware opens a server span named mcp.tool.<tool name> around each call, with the tool, the request ID and the client ID as attributes. It uses whatever tracer provider you set up, so configure the OpenTelemetry SDK to export spans over OTLP:

Pythoninventory_server.py
1
2
3
4
5
6
7
8
provider = TracerProvider(resource=Resource.create({"service.name": "inventory-mcp"}))
provider.add_span_processor(
    BatchSpanProcessor(
        OTLPSpanExporter(endpoint=os.environ.get("OTLP_TRACES_URL", "http://127.0.0.1:8412/v1/traces"))
    )
)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("inventory")
Pythoninventory_server.py
server.add_middleware(OTelMiddleware(service_name="inventory-mcp"))

Any span your tool opens becomes a child of the tool's span. That's what load_sales_history in forecast_demand is for: in a real server it would wrap the database query, so the trace shows whether the time goes into the query or somewhere else.

To see the spans without running a collector, this tiny receiver accepts OTLP over HTTP and prints one line per span. It is not an OpenTelemetry Collector: it stores nothing, forwards nothing and only understands traces. In production, point the exporter at your Collector, Jaeger, Tempo or vendor endpoint instead.

Pythonotlp_receiver.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
"""A tiny OTLP/HTTP trace receiver for local testing. Not a collector: it only prints spans."""
import json
from http.server import BaseHTTPRequestHandler, HTTPServer

from opentelemetry.proto.collector.trace.v1.trace_service_pb2 import (
    ExportTraceServiceRequest,
    ExportTraceServiceResponse,
)

STATUS = {0: "UNSET", 1: "OK", 2: "ERROR"}


def value(v):
    """Turn an OTLP AnyValue into a plain Python value."""
    return getattr(v, v.WhichOneof("value")) if v.WhichOneof("value") else None


class Receiver(BaseHTTPRequestHandler):
    def do_POST(self):
        body = self.rfile.read(int(self.headers["Content-Length"]))
        request = ExportTraceServiceRequest.FromString(body)
        for resource_spans in request.resource_spans:
            service = {a.key: value(a.value) for a in resource_spans.resource.attributes}["service.name"]
            for scope_spans in resource_spans.scope_spans:
                for span in scope_spans.spans:
                    print(
                        json.dumps(
                            {
                                "service": service,
                                "trace_id": span.trace_id.hex()[:12],
                                "span_id": span.span_id.hex()[:12],
                                "parent": span.parent_span_id.hex()[:12] or None,
                                "name": span.name,
                                "ms": round((span.end_time_unix_nano - span.start_time_unix_nano) / 1e6, 1),
                                "status": STATUS[span.status.code],
                                "attributes": {a.key: value(a.value) for a in span.attributes},
                                "events": [e.name for e in span.events],
                            }
                        ),
                        flush=True,
                    )
        payload = ExportTraceServiceResponse().SerializeToString()
        self.send_response(200)
        self.send_header("Content-Type", "application/x-protobuf")
        self.send_header("Content-Length", str(len(payload)))
        self.end_headers()
        self.wfile.write(payload)

    def log_message(self, *args):
        pass  # keep the output to spans only


if __name__ == "__main__":
    print("OTLP receiver on http://127.0.0.1:8412/v1/traces", flush=True)
    HTTPServer(("127.0.0.1", 8412), Receiver).serve_forever()

OTelMiddleware can also record a mcp.tool.duration histogram and a mcp.tool.errors counter, but only if you configure an OpenTelemetry MeterProvider. This guide uses Prometheus for metrics, so it skips that.

Step 05

Add health checks and a metrics summary

HealthCheck holds named checks, each a function that returns True or False, sync or async. A failing check with required_for_ready=True turns readiness to not_ready. MetricsCollector with MetricsMiddleware keeps a per-tool summary in memory:

Pythoninventory_server.py
health = HealthCheck()
health.add_check("warehouse_api", lambda: True, required_for_ready=True)
health.register_resources(server)
metrics.register_resource(server)

Both are MCP resources, so any MCP client can read them:

Pythonread_health.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
"""Read the health and metrics resources from the running server."""
import asyncio
import os

from promptise.mcp.client import MCPClient


async def main():
    async with MCPClient(url="http://127.0.0.1:8410/mcp", api_key=os.environ["SHOP_AGENT_KEY"]) as client:
        for uri in ("health://liveness", "health://readiness", "metrics://server"):
            result = await client.session.read_resource(uri)
            print(uri, "->", result.contents[0].text)


asyncio.run(main())

Run after the traffic in the next step, it printed:

Output
health://liveness -> {"status": "alive", "uptime_seconds": 14.0}
health://readiness -> {"status": "ready", "checks": {"warehouse_api": {"healthy": true}}}
metrics://server -> {
  "uptime_seconds": 14.0,
  "tools": {
    "check_stock": {
      "calls": 168,
      "errors": 112,
…
    "forecast_demand": {
      "calls": 15,
      "errors": 0,
      "avg_latency_ms": 1749.65
    },
…

The resource URIs are health://liveness and health://readiness. One docs page lists them as health://live and health://ready, which don't exist in 1.2.1. And because they're MCP resources, not HTTP endpoints, a Kubernetes httpGet probe can't read them; see the honest limits for what to do instead.

Step 06

Send real traffic

This script runs three copies of a well-behaved agent next to one broken bot. The agent checks stock, forecasts, reserves, and finally asks about a product that doesn't exist. The bot asks for the same stock level 150 times as fast as it can:

Pythontraffic.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
"""Drive realistic traffic at the inventory server: a normal agent and a misbehaving bot."""
import asyncio
import json
import os
from collections import Counter

from promptise.mcp.client import MCPClient

URL = "http://127.0.0.1:8410/mcp"
outcomes: Counter = Counter()


async def call(client, who, tool, args):
    result = await client.call_tool(tool, args)
    data = json.loads(result.content[0].text)
    outcome = data["error"]["code"] if isinstance(data, dict) and "error" in data else "ok"
    outcomes[(who, tool, outcome)] += 1


async def shop_agent_worker():
    """A well-behaved agent: look up, forecast, reserve."""
    async with MCPClient(url=URL, api_key=os.environ["SHOP_AGENT_KEY"]) as client:
        for _ in range(5):
            await call(client, "shop-agent", "check_stock", {"sku": "MUG-01"})
            await call(client, "shop-agent", "forecast_demand", {"sku": "MUG-01", "weeks": 2})
            await call(client, "shop-agent", "reserve_stock", {"sku": "MUG-01", "units": 1})
        await call(client, "shop-agent", "check_stock", {"sku": "NOPE-99"})


async def report_bot():
    """A bot stuck in a loop, asking for the same stock level over and over."""
    async with MCPClient(url=URL, api_key=os.environ["REPORT_BOT_KEY"]) as client:
        for _ in range(150):
            await call(client, "report-bot", "check_stock", {"sku": "DESK-07"})


async def main():
    await asyncio.gather(*(shop_agent_worker() for _ in range(3)), report_bot())
    for (who, tool, outcome), n in sorted(outcomes.items()):
        print(f"{who:11} {tool:16} {outcome:20} {n}")


asyncio.run(main())

Start the receiver, then the server, then the traffic, each in its own terminal:

>_Terminal
python otlp_receiver.py
python inventory_server.py 2> server.log
python traffic.py
Output
report-bot  check_stock      RATE_LIMIT_EXCEEDED  109
report-bot  check_stock      ok                   41
shop-agent  check_stock      TOOL_ERROR           3
shop-agent  check_stock      ok                   15
shop-agent  forecast_demand  ok                   15
shop-agent  reserve_stock    TOOL_ERROR           4
shop-agent  reserve_stock    ok                   11

The rate limiter, RateLimitMiddleware(rate_per_minute=120, burst=40), keeps one bucket per client. The bot burned through its 40-call burst and was refused 109 times. The agent's three workers stayed inside their own budget. Now find each of those problems in the signals.

Step 07

Read the metrics, logs and traces

Metrics. Scrape the exporter, here trimmed to the interesting lines:

Output
$ curl -s http://127.0.0.1:8411/metrics
# HELP mcp_tool_calls_by_client_total MCP tool calls per authenticated client, including rejected ones
# TYPE mcp_tool_calls_by_client_total counter
mcp_tool_calls_by_client_total{client="shop-agent",tool="check_stock"} 18.0
mcp_tool_calls_by_client_total{client="report-bot",tool="check_stock"} 150.0
mcp_tool_calls_by_client_total{client="shop-agent",tool="forecast_demand"} 15.0
mcp_tool_calls_by_client_total{client="shop-agent",tool="reserve_stock"} 15.0
…
# HELP mcp_tool_calls_total Total MCP tool calls
# TYPE mcp_tool_calls_total counter
mcp_tool_calls_total{status="ok",tool="check_stock"} 56.0
mcp_tool_calls_total{status="error",tool="check_stock"} 112.0
mcp_tool_calls_total{status="ok",tool="forecast_demand"} 15.0
mcp_tool_calls_total{status="error",tool="reserve_stock"} 4.0
mcp_tool_calls_total{status="ok",tool="reserve_stock"} 11.0
…
mcp_tool_duration_seconds_bucket{le="0.75",tool="forecast_demand"} 0.0
mcp_tool_duration_seconds_bucket{le="1.0",tool="forecast_demand"} 0.0
mcp_tool_duration_seconds_bucket{le="2.5",tool="forecast_demand"} 15.0
…
mcp_tool_duration_seconds_count{tool="forecast_demand"} 15.0
mcp_tool_duration_seconds_sum{tool="forecast_demand"} 26.24584025099466
…
mcp_tool_in_flight{tool="forecast_demand"} 0.0

All three problems are there. forecast_demand calls all landed between 1 and 2.5 seconds, 1.75 seconds on average (26.2 seconds over 15 calls). reserve_stock failed 4 of 15 times. And report-bot made 150 calls, eight times the agent's check_stock traffic. Notice too that the 112 check_stock errors are mostly the rate limiter's refusals: Prometheus counts them as errors like any other.

Logs. Each line in server.log is one JSON object, so jq can group failures by caller and reason:

Output
$ jq -rR 'fromjson? | select(.status == "error") | "\(.client_id)  \(.tool)  \(.error)"' server.log | sort | uniq -c
 109 report-bot  check_stock  Rate limit exceeded
   3 shop-agent  check_stock  There is no product NOPE-99.
   4 shop-agent  reserve_stock  The warehouse API did not answer. Try again in a moment.

That's the line the metrics can't give you: the bot's errors are throttling, the agent's are a bad SKU and a flaky dependency. The same file finds the slowest calls:

Output
$ jq -cnR '[inputs | fromjson? | select(.event == "tool_call_end")] | sort_by(-.duration_ms) | .[:3][] | {tool, duration_ms, request_id}' server.log
{"tool":"forecast_demand","duration_ms":2301.57,"request_id":"ed5feb201e7c"}
{"tool":"forecast_demand","duration_ms":2270.49,"request_id":"a2074245e4d2"}
{"tool":"forecast_demand","duration_ms":2180.93,"request_id":"0325c0dbed42"}

The full log lines for the slowest one:

Output
{"time": "2026-10-10T18:14:27.361+00:00", "level": "INFO", "event": "tool_call_start", "tool": "forecast_demand", "request_id": "ed5feb201e7c", "client_id": "shop-agent"}
{"time": "2026-10-10T18:14:29.662+00:00", "level": "INFO", "event": "tool_call_end", "tool": "forecast_demand", "request_id": "ed5feb201e7c", "duration_ms": 2301.57, "status": "ok", "client_id": "shop-agent"}

Traces. The request_id is also on the span, as mcp.request.id, so you can jump from a log line to its trace. The receiver printed this trace for request ed5feb201e7c:

Output
{"service": "inventory-mcp", "trace_id": "a18753e0f4bb", "span_id": "308b3ada234a", "parent": "81eb4d2e5a5c", "name": "load_sales_history", "ms": 2301.3, "status": "UNSET", "attributes": {"inventory.sku": "MUG-01"}, "events": []}
{"service": "inventory-mcp", "trace_id": "a18753e0f4bb", "span_id": "81eb4d2e5a5c", "parent": null, "name": "mcp.tool.forecast_demand", "ms": 2301.5, "status": "UNSET", "attributes": {"mcp.tool.name": "forecast_demand", "mcp.request.id": "ed5feb201e7c", "mcp.client.id": "shop-agent", "mcp.status": "ok"}, "events": []}

The child span load_sales_history took 2301.3 of the tool's 2301.5 milliseconds, so the query is the whole story. A failed call carries the error on the span:

Output
{"service": "inventory-mcp", "trace_id": "3a9e3b2c1a23", "span_id": "2a7e0d17ed9f", "parent": null, "name": "mcp.tool.reserve_stock", "ms": 0.5, "status": "ERROR", "attributes": {"mcp.tool.name": "reserve_stock", "mcp.request.id": "f3106944c37a", "mcp.client.id": "shop-agent", "mcp.status": "error", "mcp.error.message": "The warehouse API did not answer. Try again in a moment."}, "events": ["exception", "exception"]}

Successful spans keep the OpenTelemetry status UNSET, which is normal; mcp.status says ok or error. The exception is recorded twice on each failed span, once by Promptise and once by the OpenTelemetry SDK, so expect duplicate exception events in your tracing UI.

You can pick the request ID yourself. If a call arrives with an X-Request-ID header, Promptise uses it instead of generating one, which lets you follow an ID from your own system into the MCP server:

Pythonrequest_id.py
1
2
3
4
5
6
    async with MCPClient(
        url="http://127.0.0.1:8410/mcp",
        api_key=os.environ["SHOP_AGENT_KEY"],
        headers={"X-Request-ID": "order-7781"},
    ) as client:
        await client.call_tool("check_stock", {"sku": "LAMP-12"})
Output
{"time": "2026-10-10T18:14:33.996+00:00", "level": "INFO", "event": "tool_call_start", "tool": "check_stock", "request_id": "order-7781", "client_id": "shop-agent"}
{"time": "2026-10-10T18:14:33.996+00:00", "level": "INFO", "event": "tool_call_end", "tool": "check_stock", "request_id": "order-7781", "duration_ms": 0.09, "status": "ok", "client_id": "shop-agent"}
{"service": "inventory-mcp", "trace_id": "098fa3881988", "span_id": "d7a9da38f10d", "parent": null, "name": "mcp.tool.check_stock", "ms": 0.0, "status": "UNSET", "attributes": {"mcp.tool.name": "check_stock", "mcp.request.id": "order-7781", "mcp.client.id": "shop-agent", "mcp.status": "ok"}, "events": []}
Step 08

Turn the metrics into alerts

Point Prometheus at the exporter and load a rules file:

YAMLprometheus.yml
1
2
3
4
5
6
7
8
9
10
global:
  scrape_interval: 15s

rule_files:
  - mcp-alerts.yml

scrape_configs:
  - job_name: inventory-mcp
    static_configs:
      - targets: ["127.0.0.1:8411"]

Then one alert per question, built on the metric names from the scrape above:

YAMLmcp-alerts.yml
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
groups:
  - name: inventory-mcp
    rules:
      - alert: MCPToolSlow
        expr: histogram_quantile(0.95, sum by (tool, le) (rate(mcp_tool_duration_seconds_bucket[5m]))) > 2
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.tool }}: 95% of calls take up to {{ $value | humanizeDuration }}"

      - alert: MCPToolFailing
        expr: |
          sum by (tool) (rate(mcp_tool_calls_total{status="error"}[5m]))
            / sum by (tool) (rate(mcp_tool_calls_total[5m])) > 0.10
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "{{ $labels.tool }}: {{ $value | humanizePercentage }} of calls fail"

      - alert: MCPClientFlooding
        expr: sum by (client) (rate(mcp_tool_calls_by_client_total[5m])) * 60 > 100
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.client }} makes {{ $value | humanize }} tool calls a minute"

      - alert: MCPServerDown
        expr: up{job="inventory-mcp"} == 0
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Prometheus can't scrape the inventory MCP server"

For a dashboard, these queries cover the same ground: sum by (tool) (rate(mcp_tool_calls_total[5m])) for traffic per tool, topk(5, sum by (client) (rate(mcp_tool_calls_by_client_total[5m])) * 60) for the busiest callers, and sum by (tool) (mcp_tool_in_flight) for calls running right now.

Note

Docker wasn't running on the machine where this guide was written, so these rules weren't loaded into a live Prometheus. Every expression was parsed with a PromQL parser, every metric name was checked against the real scrape, and the same calculations were run over the scrape's totals. Load them with promtool check rules mcp-alerts.yml before you rely on them.

Over the whole run above, those calculations gave:

Output
  check_stock      calls= 168  error ratio=66.7%  p95≈0.005s
  forecast_demand  calls=  15  error ratio= 0.0%  p95≈2.425s
  reserve_stock    calls=  15  error ratio=26.7%  p95≈0.005s
  client=shop-agent  tool=check_stock      calls=18
  client=report-bot  tool=check_stock      calls=150
…

So MCPToolSlow would catch forecast_demand, and MCPToolFailing would catch reserve_stock. It would also catch check_stock, because the limiter's refusals count as errors. If your alert should mean "our tool is broken", move RateLimitMiddleware in front of PrometheusMiddleware, and let MCPClientFlooding cover the bot instead. Also note that the default histogram buckets jump from 1 to 2.5 seconds, so the 2.425-second p95 is interpolated, not measured. PrometheusMiddleware has no option for custom buckets in 1.2.1.

The complete server

Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
import asyncio
import json
import logging
import os
import random
import sys
from datetime import datetime, timezone

from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from prometheus_client import Counter, start_http_server

from promptise.mcp.server import (
    APIKeyAuth,
    AuthMiddleware,
    HealthCheck,
    MCPServer,
    MetricsCollector,
    MetricsMiddleware,
    OTelMiddleware,
    PrometheusMiddleware,
    RateLimitMiddleware,
    StructuredLoggingMiddleware,
    ToolError,
)

# --- Logs: one JSON object per line on stderr --------------------------------
class JsonLines(logging.Formatter):
    """Every record becomes one JSON object. Tool-call entries are JSON already."""

    def format(self, record):
        message = record.getMessage()
        try:
            entry = json.loads(message)
        except ValueError:
            entry = {"event": "log", "message": message}
        when = datetime.fromtimestamp(record.created, tz=timezone.utc).isoformat(timespec="milliseconds")
        entry = {"time": when, "level": record.levelname, **entry}
        if record.exc_info:
            entry["exception"] = self.formatException(record.exc_info)
        return json.dumps(entry)


json_lines = logging.StreamHandler(sys.stderr)
json_lines.setFormatter(JsonLines())
server_log = logging.getLogger("promptise.server")
server_log.addHandler(json_lines)
server_log.setLevel(logging.INFO)
server_log.propagate = False

# --- Traces: send every span to an OTLP endpoint -----------------------------
provider = TracerProvider(resource=Resource.create({"service.name": "inventory-mcp"}))
provider.add_span_processor(
    BatchSpanProcessor(
        OTLPSpanExporter(endpoint=os.environ.get("OTLP_TRACES_URL", "http://127.0.0.1:8412/v1/traces"))
    )
)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("inventory")

# --- The server ---------------------------------------------------------------
server = MCPServer("inventory", version="1.0.0", require_auth=True)
metrics = MetricsCollector()

CALLS_BY_CLIENT = Counter(
    "mcp_tool_calls_by_client_total",
    "MCP tool calls per authenticated client, including rejected ones",
    ["client", "tool"],
)


async def count_calls_per_client(ctx, call_next):
    """Prometheus counter with a client label, so you can see who is calling."""
    CALLS_BY_CLIENT.labels(client=ctx.client_id or "anonymous", tool=ctx.tool_name).inc()
    return await call_next(ctx)


server.add_middleware(
    AuthMiddleware(
        APIKeyAuth(
            keys={
                os.environ["SHOP_AGENT_KEY"]: {"client_id": "shop-agent", "roles": ["agent"]},
                os.environ["REPORT_BOT_KEY"]: {"client_id": "report-bot", "roles": ["agent"]},
            }
        )
    )
)
server.add_middleware(StructuredLoggingMiddleware())
server.add_middleware(PrometheusMiddleware())
server.add_middleware(OTelMiddleware(service_name="inventory-mcp"))
server.add_middleware(MetricsMiddleware(metrics))
server.add_middleware(count_calls_per_client)
server.add_middleware(RateLimitMiddleware(rate_per_minute=120, burst=40))

# A stand-in for your warehouse system.
STOCK = {"MUG-01": 42, "DESK-07": 3, "LAMP-12": 0}


@server.tool(read_only_hint=True)
async def check_stock(sku: str) -> dict:
    """How many units of a product are in the warehouse right now."""
    if sku not in STOCK:
        raise ToolError(f"There is no product {sku}.")
    return {"sku": sku, "units": STOCK[sku]}


@server.tool(read_only_hint=True)
async def forecast_demand(sku: str, weeks: int = 4) -> dict:
    """Forecast how many units of a product will sell in the next few weeks."""
    with tracer.start_as_current_span("load_sales_history") as span:
        span.set_attribute("inventory.sku", sku)
        await asyncio.sleep(random.uniform(0.8, 2.5))  # stands in for a slow query
    return {"sku": sku, "weeks": weeks, "expected_units": 6 * weeks}


@server.tool()
async def reserve_stock(sku: str, units: int) -> dict:
    """Hold units of a product for an order."""
    if random.random() < 0.25:  # stands in for a flaky warehouse API
        raise ToolError("The warehouse API did not answer. Try again in a moment.", retryable=True)
    if STOCK.get(sku, 0) < units:
        raise ToolError(f"Only {STOCK.get(sku, 0)} units of {sku} are left.")
    STOCK[sku] -= units
    return {"sku": sku, "reserved": units, "left": STOCK[sku]}


# --- Health and a metrics summary, both readable as MCP resources -------------
health = HealthCheck()
health.add_check("warehouse_api", lambda: True, required_for_ready=True)
health.register_resources(server)
metrics.register_resource(server)

if __name__ == "__main__":
    start_http_server(8411, addr="127.0.0.1")  # Prometheus scrapes http://127.0.0.1:8411/metrics
    server.run(transport="http", host="127.0.0.1", port=8410, dashboard="--dashboard" in sys.argv)

[05]

Watch it live with the dashboard

For development and incident debugging, server.run(dashboard=True) replaces the startup banner with a full-screen terminal dashboard. The server above turns it on with a flag:

>_Terminal
python inventory_server.py --dashboard

It adds its own middleware as the outermost layer and shows six tabs: Overview, Tools, Agents, Logs, Metrics and Raw Logs. Switch with the number keys, the arrow keys or Tab. This is the Metrics tab after the same traffic, captured from a 96-column terminal in a separate run, so the counts differ slightly from the run above:

Text
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
   1 Overview   2 Tools   3 Agents   4 Logs   5 Metrics   6 Raw Logs
╭────────────────────────────────────────── Metrics ───────────────────────────────────────────╮
│╭────────────────────────────────────────── Global ──────────────────────────────────────────╮│
││ 8.91 req/s             133.3ms                54.0%                 22s                    ││
││ Throughput             Avg latency            Error rate            Uptime                 ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│
│╭─────────────────────────────────── Per-Tool Performance ───────────────────────────────────╮│
││  Tool                      ┃   Calls ┃  Errors ┃    Err% ┃      Avg ┃      Min ┃      Max  ││
││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━ ││
││  check_stock               │     168 │     103 │   61.3% │    2.0ms │    0.1ms │   92.4ms  ││
││  forecast_demand           │      15 │       0 │    0.0% │ 1737.4ms │  922.7ms │ 2311.2ms  ││
││  reserve_stock             │      15 │       4 │   26.7% │    0.4ms │    0.1ms │    1.2ms  ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│
│╭───────────────────────────────────── Error Breakdown ──────────────────────────────────────╮│
││  Error Type                                                     ┃    Count ┃  % of Errors  ││
││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━ ││
││  RATE                                                           │      100 │        93.5%  ││
││  TOOL_ERROR                                                     │        7 │         6.5%  ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│

The error breakdown separates throttling from real failures, which the Prometheus counter doesn't. The Agents tab lists each client ID with its request and error counts and the tools it used, so report-bot stands out with 150 requests and 100 errors.

Know what it is before you lean on it. Everything lives in process memory and resets on restart. "Agents connected" counts every client seen since start; nobody is ever removed. While it runs, it mutes the stream handlers on promptise.server and a few other loggers and moves their output into the Raw Logs tab, so your JSON log lines stop reaching stderr. And it needs a real terminal. Use it on your laptop or in an SSH session; in production, rely on the metrics, logs and traces.


[06]

Before you go to production

These are true of Promptise Foundry 1.2.1:

  • Middleware order decides what gets counted. Calls refused by the rate limiter count as errors only for middleware added before it. Pick the order on purpose, as discussed in step 8.

  • Some failures never reach the middleware. Arguments are validated before the middleware chain runs, so a call with bad input returns a VALIDATION_ERROR to the client and appears in no metric, log line or span. A request with a wrong API key gets an HTTP 401 before MCP is involved at all; the only trace is uvicorn's access log on stdout, "POST /mcp HTTP/1.1" 401 Unauthorized. Watch your proxy or load balancer logs for both.

  • Traces start at the server. OTelMiddleware doesn't read the W3C traceparent header, so a call sent with one still starts a new trace. You can't stitch an agent's trace and the server's span into one trace yet; correlate them with X-Request-ID instead.

  • Health probes are MCP resources, not URLs. For Kubernetes, the Deployment docs suggest probing /mcp, but a plain GET /mcp gets a 406, or a 401 when authentication is on, and Kubernetes treats both as failures. Probe the metrics port, which answers 200, or use a TCP probe on the MCP port. Either only proves the process is up; readiness logic still lives in health://readiness.

  • Create `PrometheusMiddleware` once per process. A second one on the default registry raises ValueError: Duplicated timeseries in CollectorRegistry. In tests that build several servers, pass registry=CollectorRegistry().

  • Keep labels small. A client label is fine for a handful of known callers. Don't label by user, session or request ID; every new value is a new time series in Prometheus.

  • Keep an audit trail separately. Logs and metrics are for operating the server. For a tamper-evident record of who called what, add AuditMiddleware, which signs each entry with an HMAC chain. Audit logging covers its options.

  • Docs that don't match 1.2.1. The observability page shows OTelMiddleware(endpoint=...), which raises TypeError; configure the exporter on the tracer provider as in step 4. It says the Prometheus and OpenTelemetry middleware do nothing when their packages are missing; they raise ImportError. And its sample log entries differ from what StructuredLoggingMiddleware writes, shown in step 3.


[07]

Frequently asked questions

What should you monitor on an MCP server?

Per tool: call rate, error rate and latency, which PrometheusMiddleware gives you. Per caller: call counts and errors, from client_id in the logs and a small counter of your own. On top of that, a liveness and readiness check, and traces for the slow tools. Alert on latency and error rate per tool, and on any single client's call rate.

How do I add a Prometheus metrics endpoint in Python?

Install prometheus-client, record your metrics, and call start_http_server(port) once at startup; it serves /metrics from a background thread. With Promptise, PrometheusMiddleware records the tool metrics for you, and start_http_server exposes them on their own port, since the MCP server's HTTP port doesn't serve /metrics in 1.2.1.

How do I add OpenTelemetry tracing to a Python MCP server?

Install opentelemetry-sdk and an OTLP exporter, set up a TracerProvider with a BatchSpanProcessor, and add OTelMiddleware to the server. Every tool call becomes a span with the tool name, request ID and client ID. Spans you open inside a tool become its children.

What is the best way to do MCP server logging?

Log one structured JSON entry per tool call, with the tool, the caller, the duration, the status and the error text, and a request ID that also appears on your traces. StructuredLoggingMiddleware writes those fields; add a formatter for timestamps and to keep every line valid JSON. Keep arguments out of these logs, and put them in an audit log if you need them.

Does the live dashboard work in production?

It's a terminal UI that keeps its numbers in memory and needs a real terminal, so it suits development and live debugging, not a container. In production, scrape the metrics and ship the logs and traces to the tools your team already watches.


[08]

Where to go next

  • Observability & Monitoring: metrics, Prometheus, OpenTelemetry, structured and audit logging, and the dashboard.

  • Production Features: the production middleware set and the recommended order.

  • Resilience Patterns: health checks, circuit breakers and webhooks for alerting on errors.

  • Caching & Performance: rate limits and concurrency limits for the callers your metrics expose.

  • Deployment: running the server behind a proxy and in containers.

  • How to Connect MCP Servers to Your AI Agent in Python: put a real agent in front of this server.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server from an existing API, then monitor it the same way.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths