Once real agents call your MCP server, three questions come up fast: which tool is slow, which one is failing, and who is calling it far more than they should. MCP server monitoring answers them with three signals: metrics for trends and alerts, structured logs for the details of each call, and traces for where the time went. In this guide you'll build a small inventory server in Python with one slow tool and one flaky one, send it real traffic from a well-behaved agent and a misbehaving bot, and read all three signals, plus a live terminal dashboard. Every snippet was run against Promptise Foundry 1.2.1, and every output is what it printed.
How do you monitor an MCP server?
Wrap every tool call in middleware that records it, then send the records to the tools your team already uses. With Promptise Foundry, two lines of middleware give you JSON logs and Prometheus metrics, and one call from prometheus_client serves them for scraping:
import logging
from prometheus_client import start_http_server
from promptise.mcp.server import MCPServer, PrometheusMiddleware, StructuredLoggingMiddleware
logging.basicConfig(level=logging.INFO, format="%(message)s")
server = MCPServer("inventory")
server.add_middleware(StructuredLoggingMiddleware()) # one JSON log line per call start and end
server.add_middleware(PrometheusMiddleware()) # call counts, errors, latency histogram
@server.tool()
async def check_stock(sku: str) -> dict:
"""How many units of a product are in the warehouse right now."""
return {"sku": sku, "units": 42}
if __name__ == "__main__":
start_http_server(8411, addr="127.0.0.1") # Prometheus scrapes http://127.0.0.1:8411/metrics
server.run(transport="http", host="127.0.0.1", port=8410)After one tool call, the server logged two JSON lines and Prometheus could scrape the call:
{"event": "tool_call_start", "tool": "check_stock", "request_id": "e0260b6f7d1b"}
{"event": "tool_call_end", "tool": "check_stock", "request_id": "e0260b6f7d1b", "duration_ms": 0.25, "status": "ok"}$ curl -s http://127.0.0.1:8411/metrics | grep "^mcp_"
mcp_tool_calls_total{status="ok",tool="check_stock"} 1.0
…
mcp_tool_duration_seconds_count{tool="check_stock"} 1.0
mcp_tool_duration_seconds_sum{tool="check_stock"} 5.179100116947666e-05
…
mcp_tool_in_flight{tool="check_stock"} 0.0The metrics live on their own port because the MCP server's HTTP port only serves /mcp. More on that in step 2. The rest of this guide adds tracing, health checks, per-client numbers and alerts, and puts them under real load.
[02]
How it works
Every tool call passes through the server's middleware in the order you add it. Each observability middleware times the call, notes whether it raised, and hands the result to its own destination:
Rendering diagram…
Authentication goes first so every later middleware knows the caller. Each piece answers a different question:
Question | Signal | Promptise piece |
|---|---|---|
Which tool is slow? | Latency histogram, spans | PrometheusMiddleware, OTelMiddleware |
Which tool is failing, and why? | Error counter, log lines with the error text | PrometheusMiddleware, StructuredLoggingMiddleware |
Who is hammering the server? | client_id on logs and spans, a per-client counter | StructuredLoggingMiddleware, OTelMiddleware, six lines of your own |
Is it up and ready? | Liveness and readiness probes | HealthCheck |
What's happening right now? | A live terminal UI | server.run(dashboard=True) |
Who did what, provably? | A signed audit log | AuditMiddleware |
[03]
What you need
Python 3.10 or newer.
Promptise Foundry plus the Prometheus and OpenTelemetry packages. They aren't installed with plain pip install promptise, and without them PrometheusMiddleware() and OTelMiddleware() raise ImportError when you create them. pip install "promptise[all]" also pulls them in, along with everything else.
No model API key. The traffic in this guide comes from a small script, so you can rerun it as often as you like.
Optionally jq for querying the logs, and a Prometheus server and an OpenTelemetry Collector if you want to see the data in your own stack.
pip install promptise prometheus-client opentelemetry-sdk opentelemetry-exporter-otlp-proto-httpThis guide ran with prometheus-client 0.26.0 and OpenTelemetry 1.45.1.
[04]
Monitor the inventory MCP server, step by step
The server has three tools. check_stock is fast. forecast_demand is slow, between 0.8 and 2.5 seconds. reserve_stock fails about one call in four, like a flaky warehouse API. Two API keys identify the callers: shop-agent behaves, report-bot doesn't. The complete server file is at the end of this section; each step shows the part it adds.
Build the server and know who is calling
Monitoring per caller only works if the server knows the caller. AuthMiddleware with APIKeyAuth turns each key into a client_id, and because it's added first, every middleware after it can read that ID:
server = MCPServer("inventory", version="1.0.0", require_auth=True)server.add_middleware(
AuthMiddleware(
APIKeyAuth(
keys={
os.environ["SHOP_AGENT_KEY"]: {"client_id": "shop-agent", "roles": ["agent"]},
os.environ["REPORT_BOT_KEY"]: {"client_id": "report-bot", "roles": ["agent"]},
}
)
)
)The tools fail the way real tools do, with a ToolError the agent can read:
@server.tool(read_only_hint=True)
async def forecast_demand(sku: str, weeks: int = 4) -> dict:
"""Forecast how many units of a product will sell in the next few weeks."""
with tracer.start_as_current_span("load_sales_history") as span:
span.set_attribute("inventory.sku", sku)
await asyncio.sleep(random.uniform(0.8, 2.5)) # stands in for a slow query
return {"sku": sku, "weeks": weeks, "expected_units": 6 * weeks}
@server.tool()
async def reserve_stock(sku: str, units: int) -> dict:
"""Hold units of a product for an order."""
if random.random() < 0.25: # stands in for a flaky warehouse API
raise ToolError("The warehouse API did not answer. Try again in a moment.", retryable=True)
if STOCK.get(sku, 0) < units:
raise ToolError(f"Only {STOCK.get(sku, 0)} units of {sku} are left.")
STOCK[sku] -= units
return {"sku": sku, "reserved": units, "left": STOCK[sku]}The span inside forecast_demand comes into play in step 4.
Expose Prometheus metrics
PrometheusMiddleware records three metrics for every call: a counter mcp_tool_calls_total with tool and status labels, a histogram mcp_tool_duration_seconds per tool, and a gauge mcp_tool_in_flight per tool. Pass namespace="myapp" to change the mcp_ prefix.
Prometheus metrics answer "which tool", but not "which caller". A six-line middleware of your own adds a counter with a client label. It sits before the rate limiter, so calls the limiter rejects are counted too, which is exactly what you want when someone floods you:
CALLS_BY_CLIENT = Counter(
"mcp_tool_calls_by_client_total",
"MCP tool calls per authenticated client, including rejected ones",
["client", "tool"],
)
async def count_calls_per_client(ctx, call_next):
"""Prometheus counter with a client label, so you can see who is calling."""
CALLS_BY_CLIENT.labels(client=ctx.client_id or "anonymous", tool=ctx.tool_name).inc()
return await call_next(ctx)Then serve the metrics. The docstring of PrometheusMiddleware says they're available at GET /metrics on the HTTP transport, but in 1.2.1 that route doesn't exist. With a valid key, the MCP port answers:
$ curl -i -H "x-api-key: $SHOP_AGENT_KEY" http://127.0.0.1:8410/metrics
HTTP/1.1 404 Not Found
…
Not FoundSo start prometheus_client's own exporter on a second port before the server runs. It serves the default registry from a background thread:
if __name__ == "__main__":
start_http_server(8411, addr="127.0.0.1") # Prometheus scrapes http://127.0.0.1:8411/metrics
server.run(transport="http", host="127.0.0.1", port=8410, dashboard="--dashboard" in sys.argv)A separate port has an upside: the MCP port requires an API key, and Prometheus doesn't need one. Keep the metrics port on your internal network.
Log every call as JSON
StructuredLoggingMiddleware writes two entries per call to the promptise.server logger: tool_call_start, and tool_call_end with duration_ms, status, the error message on failure, and the client_id. Each entry is a JSON string, but two things stop it from being a clean JSON log on its own: the entries carry no timestamp, and Promptise writes plain-text messages to the same logger, such as MCP server running on http://127.0.0.1:8410/mcp. A small formatter fixes both, so every line your log shipper reads is one JSON object:
class JsonLines(logging.Formatter):
"""Every record becomes one JSON object. Tool-call entries are JSON already."""
def format(self, record):
message = record.getMessage()
try:
entry = json.loads(message)
except ValueError:
entry = {"event": "log", "message": message}
when = datetime.fromtimestamp(record.created, tz=timezone.utc).isoformat(timespec="milliseconds")
entry = {"time": when, "level": record.levelname, **entry}
if record.exc_info:
entry["exception"] = self.formatException(record.exc_info)
return json.dumps(entry)
json_lines = logging.StreamHandler(sys.stderr)
json_lines.setFormatter(JsonLines())
server_log = logging.getLogger("promptise.server")
server_log.addHandler(json_lines)
server_log.setLevel(logging.INFO)
server_log.propagate = FalseUvicorn keeps writing its own plain-text access log to stdout, which is useful in its own right, as you'll see under the honest limits.
Trace every call with OpenTelemetry
Metrics tell you a tool is slow. A trace tells you which part of it is slow. OTelMiddleware opens a server span named mcp.tool.<tool name> around each call, with the tool, the request ID and the client ID as attributes. It uses whatever tracer provider you set up, so configure the OpenTelemetry SDK to export spans over OTLP:
provider = TracerProvider(resource=Resource.create({"service.name": "inventory-mcp"}))
provider.add_span_processor(
BatchSpanProcessor(
OTLPSpanExporter(endpoint=os.environ.get("OTLP_TRACES_URL", "http://127.0.0.1:8412/v1/traces"))
)
)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("inventory")server.add_middleware(OTelMiddleware(service_name="inventory-mcp"))Any span your tool opens becomes a child of the tool's span. That's what load_sales_history in forecast_demand is for: in a real server it would wrap the database query, so the trace shows whether the time goes into the query or somewhere else.
To see the spans without running a collector, this tiny receiver accepts OTLP over HTTP and prints one line per span. It is not an OpenTelemetry Collector: it stores nothing, forwards nothing and only understands traces. In production, point the exporter at your Collector, Jaeger, Tempo or vendor endpoint instead.
"""A tiny OTLP/HTTP trace receiver for local testing. Not a collector: it only prints spans."""
import json
from http.server import BaseHTTPRequestHandler, HTTPServer
from opentelemetry.proto.collector.trace.v1.trace_service_pb2 import (
ExportTraceServiceRequest,
ExportTraceServiceResponse,
)
STATUS = {0: "UNSET", 1: "OK", 2: "ERROR"}
def value(v):
"""Turn an OTLP AnyValue into a plain Python value."""
return getattr(v, v.WhichOneof("value")) if v.WhichOneof("value") else None
class Receiver(BaseHTTPRequestHandler):
def do_POST(self):
body = self.rfile.read(int(self.headers["Content-Length"]))
request = ExportTraceServiceRequest.FromString(body)
for resource_spans in request.resource_spans:
service = {a.key: value(a.value) for a in resource_spans.resource.attributes}["service.name"]
for scope_spans in resource_spans.scope_spans:
for span in scope_spans.spans:
print(
json.dumps(
{
"service": service,
"trace_id": span.trace_id.hex()[:12],
"span_id": span.span_id.hex()[:12],
"parent": span.parent_span_id.hex()[:12] or None,
"name": span.name,
"ms": round((span.end_time_unix_nano - span.start_time_unix_nano) / 1e6, 1),
"status": STATUS[span.status.code],
"attributes": {a.key: value(a.value) for a in span.attributes},
"events": [e.name for e in span.events],
}
),
flush=True,
)
payload = ExportTraceServiceResponse().SerializeToString()
self.send_response(200)
self.send_header("Content-Type", "application/x-protobuf")
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
def log_message(self, *args):
pass # keep the output to spans only
if __name__ == "__main__":
print("OTLP receiver on http://127.0.0.1:8412/v1/traces", flush=True)
HTTPServer(("127.0.0.1", 8412), Receiver).serve_forever()OTelMiddleware can also record a mcp.tool.duration histogram and a mcp.tool.errors counter, but only if you configure an OpenTelemetry MeterProvider. This guide uses Prometheus for metrics, so it skips that.
Add health checks and a metrics summary
HealthCheck holds named checks, each a function that returns True or False, sync or async. A failing check with required_for_ready=True turns readiness to not_ready. MetricsCollector with MetricsMiddleware keeps a per-tool summary in memory:
health = HealthCheck()
health.add_check("warehouse_api", lambda: True, required_for_ready=True)
health.register_resources(server)
metrics.register_resource(server)Both are MCP resources, so any MCP client can read them:
"""Read the health and metrics resources from the running server."""
import asyncio
import os
from promptise.mcp.client import MCPClient
async def main():
async with MCPClient(url="http://127.0.0.1:8410/mcp", api_key=os.environ["SHOP_AGENT_KEY"]) as client:
for uri in ("health://liveness", "health://readiness", "metrics://server"):
result = await client.session.read_resource(uri)
print(uri, "->", result.contents[0].text)
asyncio.run(main())Run after the traffic in the next step, it printed:
health://liveness -> {"status": "alive", "uptime_seconds": 14.0}
health://readiness -> {"status": "ready", "checks": {"warehouse_api": {"healthy": true}}}
metrics://server -> {
"uptime_seconds": 14.0,
"tools": {
"check_stock": {
"calls": 168,
"errors": 112,
…
"forecast_demand": {
"calls": 15,
"errors": 0,
"avg_latency_ms": 1749.65
},
…The resource URIs are health://liveness and health://readiness. One docs page lists them as health://live and health://ready, which don't exist in 1.2.1. And because they're MCP resources, not HTTP endpoints, a Kubernetes httpGet probe can't read them; see the honest limits for what to do instead.
Send real traffic
This script runs three copies of a well-behaved agent next to one broken bot. The agent checks stock, forecasts, reserves, and finally asks about a product that doesn't exist. The bot asks for the same stock level 150 times as fast as it can:
"""Drive realistic traffic at the inventory server: a normal agent and a misbehaving bot."""
import asyncio
import json
import os
from collections import Counter
from promptise.mcp.client import MCPClient
URL = "http://127.0.0.1:8410/mcp"
outcomes: Counter = Counter()
async def call(client, who, tool, args):
result = await client.call_tool(tool, args)
data = json.loads(result.content[0].text)
outcome = data["error"]["code"] if isinstance(data, dict) and "error" in data else "ok"
outcomes[(who, tool, outcome)] += 1
async def shop_agent_worker():
"""A well-behaved agent: look up, forecast, reserve."""
async with MCPClient(url=URL, api_key=os.environ["SHOP_AGENT_KEY"]) as client:
for _ in range(5):
await call(client, "shop-agent", "check_stock", {"sku": "MUG-01"})
await call(client, "shop-agent", "forecast_demand", {"sku": "MUG-01", "weeks": 2})
await call(client, "shop-agent", "reserve_stock", {"sku": "MUG-01", "units": 1})
await call(client, "shop-agent", "check_stock", {"sku": "NOPE-99"})
async def report_bot():
"""A bot stuck in a loop, asking for the same stock level over and over."""
async with MCPClient(url=URL, api_key=os.environ["REPORT_BOT_KEY"]) as client:
for _ in range(150):
await call(client, "report-bot", "check_stock", {"sku": "DESK-07"})
async def main():
await asyncio.gather(*(shop_agent_worker() for _ in range(3)), report_bot())
for (who, tool, outcome), n in sorted(outcomes.items()):
print(f"{who:11} {tool:16} {outcome:20} {n}")
asyncio.run(main())Start the receiver, then the server, then the traffic, each in its own terminal:
python otlp_receiver.py
python inventory_server.py 2> server.log
python traffic.pyreport-bot check_stock RATE_LIMIT_EXCEEDED 109
report-bot check_stock ok 41
shop-agent check_stock TOOL_ERROR 3
shop-agent check_stock ok 15
shop-agent forecast_demand ok 15
shop-agent reserve_stock TOOL_ERROR 4
shop-agent reserve_stock ok 11The rate limiter, RateLimitMiddleware(rate_per_minute=120, burst=40), keeps one bucket per client. The bot burned through its 40-call burst and was refused 109 times. The agent's three workers stayed inside their own budget. Now find each of those problems in the signals.
Read the metrics, logs and traces
Metrics. Scrape the exporter, here trimmed to the interesting lines:
$ curl -s http://127.0.0.1:8411/metrics
# HELP mcp_tool_calls_by_client_total MCP tool calls per authenticated client, including rejected ones
# TYPE mcp_tool_calls_by_client_total counter
mcp_tool_calls_by_client_total{client="shop-agent",tool="check_stock"} 18.0
mcp_tool_calls_by_client_total{client="report-bot",tool="check_stock"} 150.0
mcp_tool_calls_by_client_total{client="shop-agent",tool="forecast_demand"} 15.0
mcp_tool_calls_by_client_total{client="shop-agent",tool="reserve_stock"} 15.0
…
# HELP mcp_tool_calls_total Total MCP tool calls
# TYPE mcp_tool_calls_total counter
mcp_tool_calls_total{status="ok",tool="check_stock"} 56.0
mcp_tool_calls_total{status="error",tool="check_stock"} 112.0
mcp_tool_calls_total{status="ok",tool="forecast_demand"} 15.0
mcp_tool_calls_total{status="error",tool="reserve_stock"} 4.0
mcp_tool_calls_total{status="ok",tool="reserve_stock"} 11.0
…
mcp_tool_duration_seconds_bucket{le="0.75",tool="forecast_demand"} 0.0
mcp_tool_duration_seconds_bucket{le="1.0",tool="forecast_demand"} 0.0
mcp_tool_duration_seconds_bucket{le="2.5",tool="forecast_demand"} 15.0
…
mcp_tool_duration_seconds_count{tool="forecast_demand"} 15.0
mcp_tool_duration_seconds_sum{tool="forecast_demand"} 26.24584025099466
…
mcp_tool_in_flight{tool="forecast_demand"} 0.0All three problems are there. forecast_demand calls all landed between 1 and 2.5 seconds, 1.75 seconds on average (26.2 seconds over 15 calls). reserve_stock failed 4 of 15 times. And report-bot made 150 calls, eight times the agent's check_stock traffic. Notice too that the 112 check_stock errors are mostly the rate limiter's refusals: Prometheus counts them as errors like any other.
Logs. Each line in server.log is one JSON object, so jq can group failures by caller and reason:
$ jq -rR 'fromjson? | select(.status == "error") | "\(.client_id) \(.tool) \(.error)"' server.log | sort | uniq -c
109 report-bot check_stock Rate limit exceeded
3 shop-agent check_stock There is no product NOPE-99.
4 shop-agent reserve_stock The warehouse API did not answer. Try again in a moment.That's the line the metrics can't give you: the bot's errors are throttling, the agent's are a bad SKU and a flaky dependency. The same file finds the slowest calls:
$ jq -cnR '[inputs | fromjson? | select(.event == "tool_call_end")] | sort_by(-.duration_ms) | .[:3][] | {tool, duration_ms, request_id}' server.log
{"tool":"forecast_demand","duration_ms":2301.57,"request_id":"ed5feb201e7c"}
{"tool":"forecast_demand","duration_ms":2270.49,"request_id":"a2074245e4d2"}
{"tool":"forecast_demand","duration_ms":2180.93,"request_id":"0325c0dbed42"}The full log lines for the slowest one:
{"time": "2026-10-10T18:14:27.361+00:00", "level": "INFO", "event": "tool_call_start", "tool": "forecast_demand", "request_id": "ed5feb201e7c", "client_id": "shop-agent"}
{"time": "2026-10-10T18:14:29.662+00:00", "level": "INFO", "event": "tool_call_end", "tool": "forecast_demand", "request_id": "ed5feb201e7c", "duration_ms": 2301.57, "status": "ok", "client_id": "shop-agent"}Traces. The request_id is also on the span, as mcp.request.id, so you can jump from a log line to its trace. The receiver printed this trace for request ed5feb201e7c:
{"service": "inventory-mcp", "trace_id": "a18753e0f4bb", "span_id": "308b3ada234a", "parent": "81eb4d2e5a5c", "name": "load_sales_history", "ms": 2301.3, "status": "UNSET", "attributes": {"inventory.sku": "MUG-01"}, "events": []}
{"service": "inventory-mcp", "trace_id": "a18753e0f4bb", "span_id": "81eb4d2e5a5c", "parent": null, "name": "mcp.tool.forecast_demand", "ms": 2301.5, "status": "UNSET", "attributes": {"mcp.tool.name": "forecast_demand", "mcp.request.id": "ed5feb201e7c", "mcp.client.id": "shop-agent", "mcp.status": "ok"}, "events": []}The child span load_sales_history took 2301.3 of the tool's 2301.5 milliseconds, so the query is the whole story. A failed call carries the error on the span:
{"service": "inventory-mcp", "trace_id": "3a9e3b2c1a23", "span_id": "2a7e0d17ed9f", "parent": null, "name": "mcp.tool.reserve_stock", "ms": 0.5, "status": "ERROR", "attributes": {"mcp.tool.name": "reserve_stock", "mcp.request.id": "f3106944c37a", "mcp.client.id": "shop-agent", "mcp.status": "error", "mcp.error.message": "The warehouse API did not answer. Try again in a moment."}, "events": ["exception", "exception"]}Successful spans keep the OpenTelemetry status UNSET, which is normal; mcp.status says ok or error. The exception is recorded twice on each failed span, once by Promptise and once by the OpenTelemetry SDK, so expect duplicate exception events in your tracing UI.
You can pick the request ID yourself. If a call arrives with an X-Request-ID header, Promptise uses it instead of generating one, which lets you follow an ID from your own system into the MCP server:
async with MCPClient(
url="http://127.0.0.1:8410/mcp",
api_key=os.environ["SHOP_AGENT_KEY"],
headers={"X-Request-ID": "order-7781"},
) as client:
await client.call_tool("check_stock", {"sku": "LAMP-12"}){"time": "2026-10-10T18:14:33.996+00:00", "level": "INFO", "event": "tool_call_start", "tool": "check_stock", "request_id": "order-7781", "client_id": "shop-agent"}
{"time": "2026-10-10T18:14:33.996+00:00", "level": "INFO", "event": "tool_call_end", "tool": "check_stock", "request_id": "order-7781", "duration_ms": 0.09, "status": "ok", "client_id": "shop-agent"}
{"service": "inventory-mcp", "trace_id": "098fa3881988", "span_id": "d7a9da38f10d", "parent": null, "name": "mcp.tool.check_stock", "ms": 0.0, "status": "UNSET", "attributes": {"mcp.tool.name": "check_stock", "mcp.request.id": "order-7781", "mcp.client.id": "shop-agent", "mcp.status": "ok"}, "events": []}Turn the metrics into alerts
Point Prometheus at the exporter and load a rules file:
global:
scrape_interval: 15s
rule_files:
- mcp-alerts.yml
scrape_configs:
- job_name: inventory-mcp
static_configs:
- targets: ["127.0.0.1:8411"]Then one alert per question, built on the metric names from the scrape above:
groups:
- name: inventory-mcp
rules:
- alert: MCPToolSlow
expr: histogram_quantile(0.95, sum by (tool, le) (rate(mcp_tool_duration_seconds_bucket[5m]))) > 2
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.tool }}: 95% of calls take up to {{ $value | humanizeDuration }}"
- alert: MCPToolFailing
expr: |
sum by (tool) (rate(mcp_tool_calls_total{status="error"}[5m]))
/ sum by (tool) (rate(mcp_tool_calls_total[5m])) > 0.10
for: 10m
labels:
severity: page
annotations:
summary: "{{ $labels.tool }}: {{ $value | humanizePercentage }} of calls fail"
- alert: MCPClientFlooding
expr: sum by (client) (rate(mcp_tool_calls_by_client_total[5m])) * 60 > 100
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.client }} makes {{ $value | humanize }} tool calls a minute"
- alert: MCPServerDown
expr: up{job="inventory-mcp"} == 0
for: 2m
labels:
severity: page
annotations:
summary: "Prometheus can't scrape the inventory MCP server"For a dashboard, these queries cover the same ground: sum by (tool) (rate(mcp_tool_calls_total[5m])) for traffic per tool, topk(5, sum by (client) (rate(mcp_tool_calls_by_client_total[5m])) * 60) for the busiest callers, and sum by (tool) (mcp_tool_in_flight) for calls running right now.
Over the whole run above, those calculations gave:
check_stock calls= 168 error ratio=66.7% p95≈0.005s
forecast_demand calls= 15 error ratio= 0.0% p95≈2.425s
reserve_stock calls= 15 error ratio=26.7% p95≈0.005s
client=shop-agent tool=check_stock calls=18
client=report-bot tool=check_stock calls=150
…So MCPToolSlow would catch forecast_demand, and MCPToolFailing would catch reserve_stock. It would also catch check_stock, because the limiter's refusals count as errors. If your alert should mean "our tool is broken", move RateLimitMiddleware in front of PrometheusMiddleware, and let MCPClientFlooding cover the bot instead. Also note that the default histogram buckets jump from 1 to 2.5 seconds, so the 2.425-second p95 is interpolated, not measured. PrometheusMiddleware has no option for custom buckets in 1.2.1.
The complete server
import asyncio
import json
import logging
import os
import random
import sys
from datetime import datetime, timezone
from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from prometheus_client import Counter, start_http_server
from promptise.mcp.server import (
APIKeyAuth,
AuthMiddleware,
HealthCheck,
MCPServer,
MetricsCollector,
MetricsMiddleware,
OTelMiddleware,
PrometheusMiddleware,
RateLimitMiddleware,
StructuredLoggingMiddleware,
ToolError,
)
# --- Logs: one JSON object per line on stderr --------------------------------
class JsonLines(logging.Formatter):
"""Every record becomes one JSON object. Tool-call entries are JSON already."""
def format(self, record):
message = record.getMessage()
try:
entry = json.loads(message)
except ValueError:
entry = {"event": "log", "message": message}
when = datetime.fromtimestamp(record.created, tz=timezone.utc).isoformat(timespec="milliseconds")
entry = {"time": when, "level": record.levelname, **entry}
if record.exc_info:
entry["exception"] = self.formatException(record.exc_info)
return json.dumps(entry)
json_lines = logging.StreamHandler(sys.stderr)
json_lines.setFormatter(JsonLines())
server_log = logging.getLogger("promptise.server")
server_log.addHandler(json_lines)
server_log.setLevel(logging.INFO)
server_log.propagate = False
# --- Traces: send every span to an OTLP endpoint -----------------------------
provider = TracerProvider(resource=Resource.create({"service.name": "inventory-mcp"}))
provider.add_span_processor(
BatchSpanProcessor(
OTLPSpanExporter(endpoint=os.environ.get("OTLP_TRACES_URL", "http://127.0.0.1:8412/v1/traces"))
)
)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("inventory")
# --- The server ---------------------------------------------------------------
server = MCPServer("inventory", version="1.0.0", require_auth=True)
metrics = MetricsCollector()
CALLS_BY_CLIENT = Counter(
"mcp_tool_calls_by_client_total",
"MCP tool calls per authenticated client, including rejected ones",
["client", "tool"],
)
async def count_calls_per_client(ctx, call_next):
"""Prometheus counter with a client label, so you can see who is calling."""
CALLS_BY_CLIENT.labels(client=ctx.client_id or "anonymous", tool=ctx.tool_name).inc()
return await call_next(ctx)
server.add_middleware(
AuthMiddleware(
APIKeyAuth(
keys={
os.environ["SHOP_AGENT_KEY"]: {"client_id": "shop-agent", "roles": ["agent"]},
os.environ["REPORT_BOT_KEY"]: {"client_id": "report-bot", "roles": ["agent"]},
}
)
)
)
server.add_middleware(StructuredLoggingMiddleware())
server.add_middleware(PrometheusMiddleware())
server.add_middleware(OTelMiddleware(service_name="inventory-mcp"))
server.add_middleware(MetricsMiddleware(metrics))
server.add_middleware(count_calls_per_client)
server.add_middleware(RateLimitMiddleware(rate_per_minute=120, burst=40))
# A stand-in for your warehouse system.
STOCK = {"MUG-01": 42, "DESK-07": 3, "LAMP-12": 0}
@server.tool(read_only_hint=True)
async def check_stock(sku: str) -> dict:
"""How many units of a product are in the warehouse right now."""
if sku not in STOCK:
raise ToolError(f"There is no product {sku}.")
return {"sku": sku, "units": STOCK[sku]}
@server.tool(read_only_hint=True)
async def forecast_demand(sku: str, weeks: int = 4) -> dict:
"""Forecast how many units of a product will sell in the next few weeks."""
with tracer.start_as_current_span("load_sales_history") as span:
span.set_attribute("inventory.sku", sku)
await asyncio.sleep(random.uniform(0.8, 2.5)) # stands in for a slow query
return {"sku": sku, "weeks": weeks, "expected_units": 6 * weeks}
@server.tool()
async def reserve_stock(sku: str, units: int) -> dict:
"""Hold units of a product for an order."""
if random.random() < 0.25: # stands in for a flaky warehouse API
raise ToolError("The warehouse API did not answer. Try again in a moment.", retryable=True)
if STOCK.get(sku, 0) < units:
raise ToolError(f"Only {STOCK.get(sku, 0)} units of {sku} are left.")
STOCK[sku] -= units
return {"sku": sku, "reserved": units, "left": STOCK[sku]}
# --- Health and a metrics summary, both readable as MCP resources -------------
health = HealthCheck()
health.add_check("warehouse_api", lambda: True, required_for_ready=True)
health.register_resources(server)
metrics.register_resource(server)
if __name__ == "__main__":
start_http_server(8411, addr="127.0.0.1") # Prometheus scrapes http://127.0.0.1:8411/metrics
server.run(transport="http", host="127.0.0.1", port=8410, dashboard="--dashboard" in sys.argv)[05]
Watch it live with the dashboard
For development and incident debugging, server.run(dashboard=True) replaces the startup banner with a full-screen terminal dashboard. The server above turns it on with a flag:
python inventory_server.py --dashboardIt adds its own middleware as the outermost layer and shows six tabs: Overview, Tools, Agents, Logs, Metrics and Raw Logs. Switch with the number keys, the arrow keys or Tab. This is the Metrics tab after the same traffic, captured from a 96-column terminal in a separate run, so the counts differ slightly from the run above:
1 Overview 2 Tools 3 Agents 4 Logs 5 Metrics 6 Raw Logs
╭────────────────────────────────────────── Metrics ───────────────────────────────────────────╮
│╭────────────────────────────────────────── Global ──────────────────────────────────────────╮│
││ 8.91 req/s 133.3ms 54.0% 22s ││
││ Throughput Avg latency Error rate Uptime ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│
│╭─────────────────────────────────── Per-Tool Performance ───────────────────────────────────╮│
││ Tool ┃ Calls ┃ Errors ┃ Err% ┃ Avg ┃ Min ┃ Max ││
││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━ ││
││ check_stock │ 168 │ 103 │ 61.3% │ 2.0ms │ 0.1ms │ 92.4ms ││
││ forecast_demand │ 15 │ 0 │ 0.0% │ 1737.4ms │ 922.7ms │ 2311.2ms ││
││ reserve_stock │ 15 │ 4 │ 26.7% │ 0.4ms │ 0.1ms │ 1.2ms ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│
│╭───────────────────────────────────── Error Breakdown ──────────────────────────────────────╮│
││ Error Type ┃ Count ┃ % of Errors ││
││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━ ││
││ RATE │ 100 │ 93.5% ││
││ TOOL_ERROR │ 7 │ 6.5% ││
│╰────────────────────────────────────────────────────────────────────────────────────────────╯│The error breakdown separates throttling from real failures, which the Prometheus counter doesn't. The Agents tab lists each client ID with its request and error counts and the tools it used, so report-bot stands out with 150 requests and 100 errors.
Know what it is before you lean on it. Everything lives in process memory and resets on restart. "Agents connected" counts every client seen since start; nobody is ever removed. While it runs, it mutes the stream handlers on promptise.server and a few other loggers and moves their output into the Raw Logs tab, so your JSON log lines stop reaching stderr. And it needs a real terminal. Use it on your laptop or in an SSH session; in production, rely on the metrics, logs and traces.
[06]
Before you go to production
These are true of Promptise Foundry 1.2.1:
Middleware order decides what gets counted. Calls refused by the rate limiter count as errors only for middleware added before it. Pick the order on purpose, as discussed in step 8.
Some failures never reach the middleware. Arguments are validated before the middleware chain runs, so a call with bad input returns a VALIDATION_ERROR to the client and appears in no metric, log line or span. A request with a wrong API key gets an HTTP 401 before MCP is involved at all; the only trace is uvicorn's access log on stdout,
"POST /mcp HTTP/1.1" 401 Unauthorized. Watch your proxy or load balancer logs for both.Traces start at the server. OTelMiddleware doesn't read the W3C traceparent header, so a call sent with one still starts a new trace. You can't stitch an agent's trace and the server's span into one trace yet; correlate them with X-Request-ID instead.
Health probes are MCP resources, not URLs. For Kubernetes, the Deployment docs suggest probing /mcp, but a plain GET /mcp gets a 406, or a 401 when authentication is on, and Kubernetes treats both as failures. Probe the metrics port, which answers 200, or use a TCP probe on the MCP port. Either only proves the process is up; readiness logic still lives in health://readiness.
Create `PrometheusMiddleware` once per process. A second one on the default registry raises
ValueError: Duplicated timeseries in CollectorRegistry. In tests that build several servers, pass registry=CollectorRegistry().Keep labels small. A client label is fine for a handful of known callers. Don't label by user, session or request ID; every new value is a new time series in Prometheus.
Keep an audit trail separately. Logs and metrics are for operating the server. For a tamper-evident record of who called what, add AuditMiddleware, which signs each entry with an HMAC chain. Audit logging covers its options.
Docs that don't match 1.2.1. The observability page shows OTelMiddleware(endpoint=...), which raises TypeError; configure the exporter on the tracer provider as in step 4. It says the Prometheus and OpenTelemetry middleware do nothing when their packages are missing; they raise ImportError. And its sample log entries differ from what StructuredLoggingMiddleware writes, shown in step 3.
[07]
Frequently asked questions
What should you monitor on an MCP server?
Per tool: call rate, error rate and latency, which PrometheusMiddleware gives you. Per caller: call counts and errors, from client_id in the logs and a small counter of your own. On top of that, a liveness and readiness check, and traces for the slow tools. Alert on latency and error rate per tool, and on any single client's call rate.
How do I add a Prometheus metrics endpoint in Python?
Install prometheus-client, record your metrics, and call start_http_server(port) once at startup; it serves /metrics from a background thread. With Promptise, PrometheusMiddleware records the tool metrics for you, and start_http_server exposes them on their own port, since the MCP server's HTTP port doesn't serve /metrics in 1.2.1.
How do I add OpenTelemetry tracing to a Python MCP server?
Install opentelemetry-sdk and an OTLP exporter, set up a TracerProvider with a BatchSpanProcessor, and add OTelMiddleware to the server. Every tool call becomes a span with the tool name, request ID and client ID. Spans you open inside a tool become its children.
What is the best way to do MCP server logging?
Log one structured JSON entry per tool call, with the tool, the caller, the duration, the status and the error text, and a request ID that also appears on your traces. StructuredLoggingMiddleware writes those fields; add a formatter for timestamps and to keep every line valid JSON. Keep arguments out of these logs, and put them in an audit log if you need them.
Does the live dashboard work in production?
It's a terminal UI that keeps its numbers in memory and needs a real terminal, so it suits development and live debugging, not a container. In production, scrape the metrics and ship the logs and traces to the tools your team already watches.
[08]
Where to go next
Observability & Monitoring: metrics, Prometheus, OpenTelemetry, structured and audit logging, and the dashboard.
Production Features: the production middleware set and the recommended order.
Resilience Patterns: health checks, circuit breakers and webhooks for alerting on errors.
Caching & Performance: rate limits and concurrency limits for the callers your metrics expose.
Deployment: running the server behind a proxy and in containers.
How to Connect MCP Servers to Your AI Agent in Python: put a real agent in front of this server.
OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server from an existing API, then monitor it the same way.