PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Build a Production MCP Server in Python (Step by Step)

Build an MCP server in Python ready for real agents: typed tools, clear errors, API keys with roles, rate limits, a tamper-evident audit log and tests.

Level
Intermediate
Reading time
18 min
Published
Oct 10, 2026
By
Promptise Team
  • MCP
  • MCP Server
  • Python
  • Authentication
  • Testing
  • Promptise Foundry

A first MCP server takes ten lines. A server you'd let real agents use takes a bit more thought: it has to tell the model exactly what each tool expects, refuse bad input, know who is calling, keep some callers out of some tools, survive a flood of requests, and leave a record of what happened. This guide builds a small help desk server in Python with Promptise Foundry and adds each of those, one step at a time. At the end you'll have a test suite, a real agent using the server, and an audit log you can prove nobody edited. Everything here was run against Promptise Foundry 1.2.1.

[01]

How do you build an MCP server in Python?

Install Promptise Foundry, create an MCPServer, and turn functions into tools with a decorator:

Pythonfirst_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
from promptise.mcp.server import MCPServer

server = MCPServer("helpdesk")

TICKETS = {
    "T-101": {"subject": "Login fails after a password reset", "status": "open"},
    "T-102": {"subject": "Invoice shows the wrong VAT number", "status": "open"},
}


@server.tool()
async def search_tickets(query: str) -> list[dict]:
    """Find support tickets whose subject contains the query."""
    return [
        {"ticket_id": tid, **t}
        for tid, t in TICKETS.items()
        if query.lower() in t["subject"].lower()
    ]


if __name__ == "__main__":
    server.run()

That's a working MCP server. The function name becomes the tool name, the type hints become its input schema, and the docstring becomes the description the model reads. server.run() speaks stdio, so a desktop client or an agent can start it directly. The rest of this guide is about what you add before other people's agents depend on it.


[02]

What makes an MCP server production-ready?

Here's the request path you'll end up with. Every call passes through the same checks, in order, before your code runs:

Rendering diagram…

Need

What you use

The model understands each input

Annotated types with Field descriptions and limits

Clear failures

ToolError

Related tools grouped and named

MCPRouter

Knowing who is calling

AuthMiddleware with APIKeyAuth or JWTAuth

Some tools for some callers

roles= or guards= on the tool

Protection from floods and hangs

RateLimitMiddleware, TimeoutMiddleware

A record you can trust

AuditMiddleware

Confidence before you ship

TestClient


[03]

Build the help desk server, step by step

You need Python 3.10 or newer and pip install promptise. The complete server file is at the end of this section; each step shows the part it adds.

Step 01

Describe every input for the model

The model only sees your tool's name, description and input schema, so put your knowledge there. Annotated with a pydantic Field adds a description and limits to a parameter. A docstring's Args: section works too:

Pythonhelpdesk_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
TicketId = Annotated[str, Field(description='A ticket ID, for example "T-101".', pattern=r"^T-\d+$")]


@tickets.tool(read_only_hint=True)
async def search(
    query: Annotated[str, Field(description="Words to look for in the ticket subject.", min_length=2)],
    status: Literal["open", "closed"] | None = None,
) -> list[dict]:
    """Find support tickets whose subject contains the query.

    Args:
        status: Only return tickets with this status. Leave it out to search all tickets.
    """

This is the input schema the model receives for search, generated from those few lines:

Output
{'properties': {'query': {'description': 'Words to look for in the ticket subject.', 'minLength': 2, 'title': 'Query', 'type': 'string'}, 'status': {'anyOf': [{'enum': ['open', 'closed'], 'type': 'string'}, {'type': 'null'}], 'default': None, 'description': 'Only return tickets with this status. Leave it out to search all tickets.', 'title': 'Status'}}, 'required': ['query'], 'type': 'object'}

The limits aren't just documentation. Promptise checks them before your function runs, so a malformed ticket ID never reaches your code:

Output
{
  "error": {
    "code": "VALIDATION_ERROR",
    "message": "Invalid input: ticket_id: String should match pattern '^T-\\d+$'",
    "retryable": true,
    "suggestion": "Check the parameter types and constraints in the tool description.",
    ...
Note

Parameter descriptions from Field and from Args: need Promptise Foundry 1.2.1 or newer. Earlier versions dropped them from the schema.

Step 02

Fail with a message the model can use

When something legitimately can't be done, raise ToolError with a sentence the model can pass on. It arrives as a structured error, not a crash:

Pythonhelpdesk_server.py
1
2
3
4
5
6
7
@tickets.tool(read_only_hint=True)
async def get(ticket_id: TicketId) -> dict:
    """Get one ticket with its subject, status and customer."""
    ticket = TICKETS.get(ticket_id)
    if ticket is None:
        raise ToolError(f"There is no ticket {ticket_id}.")
    return {"ticket_id": ticket_id, **ticket}
Output
{
  "error": {
    "code": "TOOL_ERROR",
    "message": "There is no ticket T-999.",
    "retryable": false
  }
}
Step 03

Group related tools with a router

As a server grows, routers keep it organised. A router's prefix becomes part of each tool name, so search, get and close on the tickets router reach the model as tickets_search, tickets_get and tickets_close:

Pythonhelpdesk_server.py
tickets = MCPRouter(prefix="tickets")

# ... tools defined with @tickets.tool() ...

server.include_router(tickets)
Step 04

Know who is calling, and limit what they can do

Turn on authentication for the whole server with require_auth=True, and give each caller an API key. The rich key format attaches roles to a key:

Pythonhelpdesk_server.py
1
2
3
4
5
6
7
8
9
10
11
12
server = MCPServer("helpdesk", version="1.0.0", require_auth=True)

server.add_middleware(
    AuthMiddleware(
        APIKeyAuth(
            keys={
                os.environ["HELPDESK_AGENT_KEY"]: {"client_id": "support-bot", "roles": ["agent"]},
                os.environ["HELPDESK_LEAD_KEY"]: {"client_id": "team-lead", "roles": ["agent", "lead"]},
            }
        )
    )
)

Then decide, per tool, who may call it. Closing a ticket is for team leads only:

Pythonhelpdesk_server.py
1
2
3
4
5
6
@tickets.tool(roles=["lead"], destructive_hint=True)
async def close(
    ticket_id: TicketId,
    resolution: Annotated[str, Field(description="One sentence the customer will see.", min_length=10)],
) -> dict:
    """Close a ticket with a resolution. Only team leads can close tickets."""

When the support bot tries anyway, the answer says exactly why:

Output
{
  "error": {
    "code": "ACCESS_DENIED",
    "message": "Requires any of roles [lead], but client has [agent]",
    "retryable": false,
    "details": {
      "guard": "HasRole",
      "tool": "tickets_close"
    }
  }
}

read_only_hint and destructive_hint are standard MCP tool annotations. They tell clients what kind of tool this is, so a client can, for example, ask the user before running a destructive one. If you're using JSON Web Tokens from an identity provider instead of API keys, swap APIKeyAuth for JWTAuth. Authentication & Security covers that, plus scope-based guards.

Step 05

Add rate limits, timeouts and an audit log

Middleware wraps every tool call. It runs in the order you add it, so the first one wraps all the others:

Pythonhelpdesk_server.py
server.add_middleware(LoggingMiddleware())
server.add_middleware(AuthMiddleware(...))  # from step 4
server.add_middleware(RateLimitMiddleware(rate_per_minute=120, per_tool=True))
server.add_middleware(TimeoutMiddleware(default_timeout=10))
server.add_middleware(AuditMiddleware("audit.jsonl", hmac_secret=os.environ["HELPDESK_AUDIT_SECRET"]))
  • RateLimitMiddleware caps how often tools can be called, here 120 calls a minute per tool, so one runaway agent can't flood your backend.

  • TimeoutMiddleware stops any call that takes longer than ten seconds and returns a TIMEOUT error the agent can retry.

  • AuditMiddleware writes one line per call: which tool, which client, which roles, and whether it worked.

Here are the three audit lines from a real run, shortened:

Output
{"timestamp": 1791649006.041137, "tool": "tickets_search", "client_id": "support-bot", "request_id": "4bafb147007d", "status": "ok", …
{"timestamp": 1791649006.906142, "tool": "tickets_get", "client_id": "support-bot", "request_id": "ae00916e1345", "status": "ok", …
{"timestamp": 1791649011.448718, "tool": "tickets_close", "client_id": "support-bot", "request_id": "ccd45ce25d8e", "status": "error", …

Each line also carries the signature of the line before it, signed with your secret. That makes the log tamper-evident: change a field, delete a line or reorder two, and the chain breaks. This short script checks a log file:

Pythonverify_audit.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
import hashlib
import hmac
import json
import os
import sys


def verify(path: str, secret: str) -> bool:
    """Check that no line in an AuditMiddleware log was edited, removed or reordered."""
    prev = "0" * 64
    with open(path) as f:
        for number, line in enumerate(f, 1):
            entry = json.loads(line)
            signature = entry.pop("hmac")
            payload = json.dumps(entry, sort_keys=True, default=str).encode()
            expected = hmac.new(secret.encode(), payload, hashlib.sha256).hexdigest()
            if entry["prev_hash"] != prev or not hmac.compare_digest(signature, expected):
                print(f"line {number}: chain broken")
                return False
            prev = signature
    return True


if __name__ == "__main__":
    ok = verify(sys.argv[1], os.environ["HELPDESK_AUDIT_SECRET"])
    print("audit log intact" if ok else "audit log was changed")
    sys.exit(0 if ok else 1)

On the untouched log it prints audit log intact. After changing a single status from ok to error on the first line, it prints line 1: chain broken. It does the same for a deleted line, or the wrong secret.

Warning

Always pass hmac_secret, and keep it somewhere safe. Without one, Promptise generates a secret for the process, and a log signed with it can't be verified once the process has stopped.

Calls rejected before the audit middleware, such as those without a valid key or with invalid input, don't reach the log. Move AuditMiddleware earlier in the chain if you want those recorded too.

Step 06

Test it without starting anything

TestClient runs the full pipeline in memory: validation, authentication, middleware, role checks and your code. Pass request metadata to act as a particular caller:

Pythontest_helpdesk.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
import os

os.environ.setdefault("HELPDESK_AGENT_KEY", "test-agent-key")
os.environ.setdefault("HELPDESK_LEAD_KEY", "test-lead-key")
os.environ.setdefault("HELPDESK_AUDIT_SECRET", "test-audit-secret")

from promptise.mcp.server import TestClient  # noqa: E402

from helpdesk_server import server  # noqa: E402

AGENT = TestClient(server, meta={"x-api-key": "test-agent-key"})
LEAD = TestClient(server, meta={"x-api-key": "test-lead-key"})


async def test_tools_are_listed_with_descriptions():
    tools = {t.name: t for t in await AGENT.list_tools()}
    assert set(tools) == {"tickets_search", "tickets_get", "tickets_close"}
    status = tools["tickets_search"].inputSchema["properties"]["status"]
    assert "Only return tickets with this status" in status["description"]


async def test_search_finds_open_login_tickets():
    result = await AGENT.call_tool("tickets_search", {"query": "login", "status": "open"})
    assert "T-101" in result[0].text and "T-103" not in result[0].text


async def test_an_agent_cannot_close_tickets():
    result = await AGENT.call_tool("tickets_close", {"ticket_id": "T-102", "resolution": "Corrected the VAT number."})
    assert "ACCESS_DENIED" in result[0].text


async def test_a_lead_can_close_tickets():
    result = await LEAD.call_tool("tickets_close", {"ticket_id": "T-102", "resolution": "Corrected the VAT number."})
    assert '"status": "closed"' in result[0].text


async def test_bad_input_is_rejected_before_the_handler_runs():
    result = await AGENT.call_tool("tickets_get", {"ticket_id": "101"})
    assert "VALIDATION_ERROR" in result[0].text


async def test_calls_without_a_key_are_refused():
    result = await TestClient(server).call_tool("tickets_get", {"ticket_id": "T-101"})
    assert "AUTHENTICATION" in result[0].text.upper()

Install the test tools and run them. The tests are async, so tell pytest-asyncio to run them automatically:

>_Terminal
pip install pytest pytest-asyncio
pytest -o asyncio_mode=auto
Output
......                                                                   [100%]
6 passed in 0.80s

These tests caught a real mistake while this guide was being written: the first version of a ticket said "log in", and a search for "login" quietly found nothing.

The complete server

Pythonhelpdesk_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
import os
from typing import Annotated, Literal

from pydantic import Field

from promptise.mcp.server import (
    APIKeyAuth,
    AuditMiddleware,
    AuthMiddleware,
    LoggingMiddleware,
    MCPRouter,
    MCPServer,
    RateLimitMiddleware,
    TimeoutMiddleware,
    ToolError,
)

server = MCPServer("helpdesk", version="1.0.0", require_auth=True)

# Middleware runs in the order it's added: the first one wraps all the others.
server.add_middleware(LoggingMiddleware())
server.add_middleware(
    AuthMiddleware(
        APIKeyAuth(
            keys={
                os.environ["HELPDESK_AGENT_KEY"]: {"client_id": "support-bot", "roles": ["agent"]},
                os.environ["HELPDESK_LEAD_KEY"]: {"client_id": "team-lead", "roles": ["agent", "lead"]},
            }
        )
    )
)
server.add_middleware(RateLimitMiddleware(rate_per_minute=120, per_tool=True))
server.add_middleware(TimeoutMiddleware(default_timeout=10))
server.add_middleware(AuditMiddleware("audit.jsonl", hmac_secret=os.environ["HELPDESK_AUDIT_SECRET"]))

# A stand-in for your ticketing system.
TICKETS = {
    "T-101": {"subject": "Login fails after a password reset", "status": "open", "customer": "acme"},
    "T-102": {"subject": "Invoice shows the wrong VAT number", "status": "open", "customer": "globex"},
    "T-103": {"subject": "Login page is slow", "status": "closed", "customer": "acme"},
}

TicketId = Annotated[str, Field(description='A ticket ID, for example "T-101".', pattern=r"^T-\d+$")]

tickets = MCPRouter(prefix="tickets")


@tickets.tool(read_only_hint=True)
async def search(
    query: Annotated[str, Field(description="Words to look for in the ticket subject.", min_length=2)],
    status: Literal["open", "closed"] | None = None,
) -> list[dict]:
    """Find support tickets whose subject contains the query.

    Args:
        status: Only return tickets with this status. Leave it out to search all tickets.
    """
    words = query.lower()
    return [
        {"ticket_id": tid, **t}
        for tid, t in TICKETS.items()
        if words in t["subject"].lower() and (status is None or t["status"] == status)
    ]


@tickets.tool(read_only_hint=True)
async def get(ticket_id: TicketId) -> dict:
    """Get one ticket with its subject, status and customer."""
    ticket = TICKETS.get(ticket_id)
    if ticket is None:
        raise ToolError(f"There is no ticket {ticket_id}.")
    return {"ticket_id": ticket_id, **ticket}


@tickets.tool(roles=["lead"], destructive_hint=True)
async def close(
    ticket_id: TicketId,
    resolution: Annotated[str, Field(description="One sentence the customer will see.", min_length=10)],
) -> dict:
    """Close a ticket with a resolution. Only team leads can close tickets."""
    ticket = TICKETS.get(ticket_id)
    if ticket is None:
        raise ToolError(f"There is no ticket {ticket_id}.")
    ticket.update(status="closed", resolution=resolution)
    return {"ticket_id": ticket_id, **ticket}


server.include_router(tickets)

if __name__ == "__main__":
    server.run(transport="http", host="127.0.0.1", port=8080)

[04]

Put a real agent in front of it

Start the server with the three secrets set, and it listens at http://127.0.0.1:8080/mcp. Then connect an agent with the support bot's key. This is the same pattern as in How to Connect MCP Servers to Your AI Agent in Python:

Pythonsupport_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
import asyncio
import os

from promptise import build_agent
from promptise.config import HTTPServerSpec


async def main():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "helpdesk": HTTPServerSpec(
                url="http://127.0.0.1:8080/mcp",
                api_key=os.environ["HELPDESK_AGENT_KEY"],
            ),
        },
        instructions="You are a support assistant. Use the help desk tools.",
        trace_tools=True,
    )
    try:
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Find the open login ticket and close it as fixed."}]}
        )
        print(">>>", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
→ Invoking tool: tickets_search with {'query': 'login', 'status': 'open'}
✔ Tool result from tickets_search: [{"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open", "customer": "acme"}]
→ Invoking tool: tickets_get with {'ticket_id': 'T-101'}
✔ Tool result from tickets_get: {"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open", "customer": "acme"}
→ Invoking tool: tickets_close with {'ticket_id': 'T-101', 'resolution': 'Fixed.'}
✔ Tool result from tickets_close: Input validation error: 'Fixed.' is too short
→ Invoking tool: tickets_close with {'ticket_id': 'T-101', 'resolution': 'Fixed: corrected the password reset workflow so users can log in successfully after resetting their password.'}
✔ Tool result from tickets_close: {
  "error": {
    "code": "ACCESS_DENIED",
    "message": "Requires any of roles [lead], but client has [agent]",
…
>>> I found the open login ticket but could not close it because I don’t have the required lead role. …

Read it from the top and every rule from this guide shows up. The agent found the ticket. Its first resolution, "Fixed.", was too short for min_length=10, so it wrote a proper one. Then the role check stopped it, because the support bot isn't a lead, and the agent told the user why. Nothing in the agent's code mentions any of these rules; they live in the server, where they hold no matter which agent calls.


[05]

Before you deploy

  • Keep secrets out of code. API keys and the audit secret come from environment variables or a secret manager.

  • Put HTTPS in front. Run the server behind a reverse proxy or load balancer that terminates TLS. When you bind to a public address, pass allowed_hosts to server.run() with the hostname your proxy forwards.

  • Use your identity provider in production. API keys suit service-to-service calls. For people and per-user access, use JWTAuth, and see Multi-Tenancy if several customers share one server.

  • Ask a human before risky actions. Mark a tool with requires_approval=True and connect an approver, and each call waits for a person's sign-off. Approval Gates shows how to set it up.


[06]

Frequently asked questions

Do I need async functions for tools?

The examples use async def, which suits tools that call databases or other APIs without blocking. Keep slow, blocking work out of your tools or move it to a background job, so one call doesn't hold up the others.

stdio or HTTP?

Use stdio while you build and for tools that run on the same machine as the client. Use HTTP for anything shared or deployed. Authentication and per-caller roles only make sense over HTTP, where each request carries its own key or token.

How do I see what my server exposes?

Connect any MCP client and list its tools, or call list_tools() on a TestClient as the tests above do. Each tool's description and input schema are exactly what the model will read.

Can I use JWTs instead of API keys?

Yes. JWTAuth validates tokens, puts their roles and scopes on the request, and works with the same roles= and guards. Authentication & Security has the details.


[07]

Where to go next

  • Server fundamentals: tools, resources, prompts and dependency injection.

  • Routers & Middleware: every built-in middleware and how to write your own.

  • Testing: more on TestClient.

  • Production Features and Deployment: caching, health, queues and running in production.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server like this from an existing API.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths