PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

MCP Transports: stdio vs Streamable HTTP vs SSE (Python)

stdio vs Streamable HTTP vs SSE, shown on the wire with curl, then an MCP server over HTTP in production: Host checks, auth, health checks and scaling.

Level
Intermediate
Reading time
22 min
Published
Oct 10, 2026
By
Promptise Team
  • MCP
  • Streamable HTTP
  • MCP Transports
  • Deployment
  • Python
  • Promptise Foundry

Every MCP server talks to its clients over a transport, and the choice decides more than it looks: who starts the server, how many agents can share it, how you secure it, and how you deploy it. MCP has three: stdio, Streamable HTTP, and the older HTTP with SSE. This guide compares them on the wire with curl, connects one Python agent over all three, then takes the HTTP version to production with Host checks, API keys, a health check and a plan for restarts. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed.

[01]

MCP stdio vs Streamable HTTP vs SSE: which transport should you use?

Use stdio when the client should start the server itself on the same machine, as Claude Desktop, IDEs and local agents do. Use Streamable HTTP when the server runs on its own and several agents connect to it by URL; that's the transport for anything shared, remote or behind authentication. Use SSE only for a client that can't speak Streamable HTTP yet, because the MCP specification replaced it with Streamable HTTP in 2025.

In Promptise Foundry the same server code runs on all three. Only the last line changes:

Pythoninventory_server.py
    server.run(transport=transport, host=os.environ.get("HOST", "127.0.0.1"), port=port)

transport is "stdio", "http" or "sse". On the agent side you pick the matching spec:

Pythonagent.py
1
2
3
4
5
6
7
8
SPECS = {
    # The agent starts the server as a child process and talks over its stdin/stdout.
    "stdio": StdioServerSpec(command=sys.executable, args=["inventory_server.py"]),
    # The server is already running; the agent connects to its /mcp endpoint.
    "http": HTTPServerSpec(url="http://127.0.0.1:8170/mcp"),
    # The legacy transport: a /sse stream plus a /messages/ endpoint.
    "sse": HTTPServerSpec(url="http://127.0.0.1:8171/sse", transport="sse"),
}

If you've only ever switched server.run() to HTTP, as in How to Connect MCP Servers to Your AI Agent in Python, this guide shows what actually changes underneath, and what that means once real traffic arrives.


[02]

How the three MCP transports work

All three carry the same JSON-RPC messages. What differs is the pipe:

Rendering diagram…

  • stdio: the client launches the server and writes one JSON message per line to its standard input. Answers come back on standard output. There's no network and no port. When the client closes the pipe, the server exits.

  • Streamable HTTP: the server listens on a single endpoint, /mcp. Every message is its own POST. The server answers each request with an SSE stream scoped to that request, or plain JSON, and tracks the conversation with an mcp-session-id header.

  • SSE: the client opens a long-lived GET /sse stream first. The server tells it where to post messages, and every answer arrives on that one stream, not on the POST that asked for it.

Streamable HTTP was introduced in MCP protocol version 2025-03-26 as the replacement for HTTP with SSE from 2024-11-05, as the 2025-03-26 changelog records. Claude Code's MCP docs now call SSE deprecated and recommend HTTP for remote servers.

stdio

Streamable HTTP

SSE (legacy)

Who starts the server

The client, as a child process

You do; it runs on its own

You do; it runs on its own

On the wire

One JSON-RPC message per line on stdin and stdout

Each message is a POST /mcp; replies are an SSE stream or JSON

Replies on a long-lived GET /sse stream; requests to POST /messages/

Sessions

One per process

mcp-session-id header, held in server memory

session_id in the message URL, held in server memory

Auth

None on the wire; the server runs as the user

API key or bearer token on every request

API key or bearer token on every request

Scaling

One client per process

Replicas behind a load balancer with sticky sessions

Same as HTTP, plus one open stream per client

Clients

Claude Desktop, Claude Code, Cursor, Promptise

Claude Code (recommended), Cursor, Promptise

Claude Code (deprecated there), Cursor, Promptise

Promptise server

server.run()

server.run(transport="http")

server.run(transport="sse")

Promptise client

StdioServerSpec

HTTPServerSpec(url=".../mcp")

HTTPServerSpec(url=".../sse", transport="sse")

Client support is taken from the Claude Code and Cursor docs.

Note

The newest MCP revision, 2026-07-28, removes protocol-level sessions, the mcp-session-id header and the initialize handshake from Streamable HTTP (changelog). Promptise Foundry 1.2.1 ships with MCP SDK 1.30.0, which speaks protocol versions up to 2025-11-25, so its Streamable HTTP still uses sessions. Everything in this guide describes that behaviour, which is what clients use today.


[03]

What you need

  • Python 3.10 or newer, and pip install "promptise==1.2.1". The guide used a fresh virtual environment with nothing else installed.

  • curl, to look at the raw traffic.

  • An OpenAI API key for the one agent step. Every other step runs without a model.


[04]

See each MCP transport on the wire, step by step

Step 01

Write one server that runs on any transport

A transport is a deployment decision, not a code decision, so keep it out of your tools. This server reads the transport from the command line:

Pythoninventory_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
import os
import sys

from promptise.mcp.server import MCPServer

server = MCPServer("inventory", version="1.0.0")

# A stand-in for your warehouse system.
STOCK = {"SKU-1": 42, "SKU-2": 0, "SKU-3": 7}


@server.tool(read_only_hint=True)
async def check_stock(sku: str) -> dict:
    """Get how many units of a product are in stock.

    Args:
        sku: The product code, for example "SKU-1".
    """
    sku = sku.strip().upper()
    if sku not in STOCK:
        return {"error": f"Unknown product {sku}."}
    return {"sku": sku, "in_stock": STOCK[sku]}


if __name__ == "__main__":
    # python inventory_server.py              -> stdio
    # python inventory_server.py http 8170    -> Streamable HTTP on /mcp
    # python inventory_server.py sse 8171     -> legacy SSE on /sse
    transport = sys.argv[1] if len(sys.argv) > 1 else "stdio"
    port = int(sys.argv[2]) if len(sys.argv) > 2 else 8080
    server.run(transport=transport, host=os.environ.get("HOST", "127.0.0.1"), port=port)

Notice the explicit host. server.run() binds 0.0.0.0 by default, which listens on every network interface. Step 1 of the production section shows why 127.0.0.1 is the safer default.

Step 02

Speak stdio by hand

stdio is the simplest transport, and you can drive it with nothing but a pipe. Send the handshake, then a tool call, one JSON object per line:

>_Terminalstdio_raw.sh
# Talk to the stdio server by hand: three JSON-RPC messages, one per line.
printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"by-hand","version":"0"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"check_stock","arguments":{"sku":"SKU-1"}}}' \
  | python inventory_server.py
Output
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-11-25","capabilities":{"experimental":{},"prompts":{"listChanged":false},"resources":{"subscribe":false,"listChanged":false},"tools":{"listChanged":false}},"serverInfo":{"name":"inventory","version":"1.0.0"}}}
{"jsonrpc":"2.0","id":2,"result":{"content":[{"type":"text","text":"{\"sku\": \"SKU-1\", \"in_stock\": 42}"}],"isError":false}}

Two answers, one per line, and then the server exited because its input closed. That's the whole stdio lifecycle: the client owns the process. It also explains the one stdio rule that bites people: stdout belongs to the protocol. A stray print() in your server corrupts the stream. Log to stderr instead. Promptise skips its startup banner on stdio for this reason.

Step 03

Open a Streamable HTTP session with curl

Start the same file over HTTP with python inventory_server.py http 8170. It now listens at http://127.0.0.1:8170/mcp. A client has to accept both JSON and an event stream, so send both in Accept:

>_Terminalhttp_raw.sh
URL=http://127.0.0.1:8170/mcp
ACCEPT="Accept: application/json, text/event-stream"
JSON="Content-Type: application/json"

echo "### 1. initialize"
curl -si $URL -H "$ACCEPT" -H "$JSON" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}' \
  | tee init.txt
SESSION=$(grep -i '^mcp-session-id:' init.txt | cut -d' ' -f2 | tr -d '\r')

echo; echo "### 2. initialized notification"
curl -si $URL -H "$ACCEPT" -H "$JSON" -H "mcp-session-id: $SESSION" \
  -d '{"jsonrpc":"2.0","method":"notifications/initialized"}'

echo; echo "### 3. tools/call in the session"
curl -s $URL -H "$ACCEPT" -H "$JSON" -H "mcp-session-id: $SESSION" \
  -d '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"check_stock","arguments":{"sku":"SKU-3"}}}'
Output
### 1. initialize
HTTP/1.1 200 OK
date: Sat, 10 Oct 2026 16:25:37 GMT
server: uvicorn
cache-control: no-cache, no-transform
connection: keep-alive
content-type: text/event-stream
mcp-session-id: eae679ada7644f6088c2c4099a233781
x-accel-buffering: no
transfer-encoding: chunked

event: message
data: {"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-11-25","capabilities":{"experimental":{},"prompts":{"listChanged":false},"resources":{"subscribe":false,"listChanged":false},"tools":{"listChanged":false}},"serverInfo":{"name":"inventory","version":"1.0.0"}}}


### 2. initialized notification
HTTP/1.1 202 Accepted
…
### 3. tools/call in the session
event: message
data: {"jsonrpc":"2.0","id":2,"result":{"content":[{"type":"text","text":"{\"sku\": \"SKU-3\", \"in_stock\": 7}"}],"isError":false}}

Three things to notice:

  • The session ID arrives as a response header. The client must send mcp-session-id back on every later request. That's the "session" in Streamable HTTP.

  • Even a single answer comes back as `text/event-stream`. Promptise runs the SDK's session manager with streaming responses, so a reverse proxy must not buffer them. The x-accel-buffering: no header asks Nginx not to.

  • A notification gets `202 Accepted` with an empty body. Only requests get answers.

Step 04

Learn the session rules

Sessions are where Streamable HTTP surprises people, so try the edges. The same script continues:

>_Terminalhttp_raw.sh
echo; echo "### 4. the same call without the session header"
curl -si $URL -H "$ACCEPT" -H "$JSON" \
  -d '{"jsonrpc":"2.0","id":3,"method":"tools/list"}'
Output
### 4. the same call without the session header
HTTP/1.1 400 Bad Request
…
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Bad Request: Missing session ID"}}
…
### 7. end the session
HTTP/1.1 200 OK
…
### 8. use the ended session
HTTP/1.1 404 Not Found
…
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Session not found"}}

Steps 7 and 8 sent DELETE /mcp with the session header, then tried to keep using it. So:

  • Forgot the header? 400 Missing session ID. If you see this from a real client, a proxy is probably stripping the mcp-session-id header.

  • Session gone? 404 Session not found. The specification says a client must then start a new session with a fresh initialize. Keep that rule in mind; it matters again in production.

  • Done? DELETE ends the session cleanly, and Promptise's client sends it for you on shutdown.

Step 05

Try the legacy SSE transport

Start the server with python inventory_server.py sse 8171, open the stream in the background, and post to the address it gives you:

>_Terminalsse_raw.sh
BASE=http://127.0.0.1:8171
JSON="Content-Type: application/json"

# 1. Open the event stream and keep it open in the background.
curl -sN $BASE/sse > stream.txt &
STREAM=$!
sleep 1
echo "### the stream's first event"
cat stream.txt
ENDPOINT=$(grep '^data:' stream.txt | head -1 | cut -d' ' -f2 | tr -d '\r')

# 2. Send messages to the endpoint the server named. Each POST only says "Accepted".
echo "### POST initialize"
curl -si "$BASE$ENDPOINT" -H "$JSON" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'
echo
curl -s "$BASE$ENDPOINT" -H "$JSON" -d '{"jsonrpc":"2.0","method":"notifications/initialized"}'; echo
curl -s "$BASE$ENDPOINT" -H "$JSON" \
  -d '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"check_stock","arguments":{"sku":"SKU-2"}}}'; echo
sleep 1

# 3. The answers arrived on the stream, not on the POSTs.
echo; echo "### the stream after the POSTs"
cat stream.txt
kill $STREAM
Output
### the stream's first event
event: endpoint
data: /messages/?session_id=6ac7c6587f9747c8b00b48229deb8d3e

### POST initialize
HTTP/1.1 202 Accepted
date: Sat, 10 Oct 2026 16:25:58 GMT
server: uvicorn
content-length: 8

Accepted
Accepted
Accepted

### the stream after the POSTs
event: endpoint
data: /messages/?session_id=6ac7c6587f9747c8b00b48229deb8d3e

event: message
data: {"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2024-11-05",…

event: message
data: {"jsonrpc":"2.0","id":2,"result":{"content":[{"type":"text","text":"{\"sku\": \"SKU-2\", \"in_stock\": 0}"}],"isError":false}}

This is why SSE was replaced. Every POST gets a bare 202 Accepted, and the real answers travel down one stream that must stay open for the whole conversation. Each client holds a connection open, the request and its answer take different paths through your proxies, and a dropped stream drops everything in flight. Streamable HTTP keeps one endpoint and lets each request carry its own answer.

Step 06

Connect an agent over each transport

Now let an agent use all three. This is the full script whose SPECS you saw at the top:

Pythonagent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
import asyncio
import sys

from promptise import build_agent
from promptise.config import HTTPServerSpec, StdioServerSpec

SPECS = {
    # The agent starts the server as a child process and talks over its stdin/stdout.
    "stdio": StdioServerSpec(command=sys.executable, args=["inventory_server.py"]),
    # The server is already running; the agent connects to its /mcp endpoint.
    "http": HTTPServerSpec(url="http://127.0.0.1:8170/mcp"),
    # The legacy transport: a /sse stream plus a /messages/ endpoint.
    "sse": HTTPServerSpec(url="http://127.0.0.1:8171/sse", transport="sse"),
}


async def main(transport: str):
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"inventory": SPECS[transport]},
        instructions="You answer stock questions. Use your tools. Reply in one sentence.",
        trace_tools=True,
    )
    try:
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Do we have SKU-3 in stock?"}]}
        )
        print(f"[{transport}]", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main(sys.argv[1]))
Output
$ python agent.py stdio
→ Invoking tool: check_stock with {'sku': 'SKU-3'}
✔ Tool result from check_stock: {"sku": "SKU-3", "in_stock": 7}
[stdio] Yes — we have 7 units of SKU-3 in stock.
$ python agent.py http
→ Invoking tool: check_stock with {'sku': 'SKU-3'}
✔ Tool result from check_stock: {"sku": "SKU-3", "in_stock": 7}
[http] Yes — there are 7 units of SKU-3 in stock.
$ python agent.py sse
→ Invoking tool: check_stock with {'sku': 'SKU-3'}
✔ Tool result from check_stock: {"sku": "SKU-3", "in_stock": 7}
[sse] Yes — we have 7 units of SKU-3 in stock.

Same tool, same arguments, same answer. Above the transport, nothing changes, which is exactly why you can start on stdio and move to HTTP later.

HTTPServerSpec accepts transport="http", "streamable-http" or "sse". In Promptise 1.2.1, "http" and "streamable-http" are two names for the same Streamable HTTP client; the source routes both to the SDK's streamablehttp_client. If you don't need an agent, MCPClient takes the same transport names. This script connects to both running servers and starts a stdio copy:

Pythonthree_clients.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import asyncio
import sys

from promptise.mcp.client import MCPClient

CLIENTS = {
    "stdio": MCPClient(transport="stdio", command=sys.executable, args=["inventory_server.py"]),
    "http": MCPClient(url="http://127.0.0.1:8170/mcp", transport="http"),
    "streamable-http": MCPClient(url="http://127.0.0.1:8170/mcp", transport="streamable-http"),
    "sse": MCPClient(url="http://127.0.0.1:8171/sse", transport="sse"),
}


async def main():
    for name, client in CLIENTS.items():
        async with client:
            tools = [t.name for t in await client.list_tools()]
            result = await client.call_tool("check_stock", {"sku": "SKU-1"})
            print(f"{name:16} tools={tools}  {result.content[0].text}")


asyncio.run(main())
Output
stdio            tools=['check_stock']  {"sku": "SKU-1", "in_stock": 42}
http             tools=['check_stock']  {"sku": "SKU-1", "in_stock": 42}
streamable-http  tools=['check_stock']  {"sku": "SKU-1", "in_stock": 42}
sse              tools=['check_stock']  {"sku": "SKU-1", "in_stock": 42}

When the client and server disagree about the transport, build_agent stops before the model is ever called. These are the errors from pointing each spec at the wrong place:

Output
Streamable HTTP client, SSE server: MCPConnectionRejectedError: Server 'inventory' rejected the connection: 405 Method Not Allowed.
SSE client, Streamable HTTP server: MCPConnectionRejectedError: Server 'inventory' rejected the connection: 400 Bad Request.
URL without /mcp: MCPClientError: Failed to connect to server 'inventory': Failed to connect to http://127.0.0.1:8170 (McpError: Session terminated)

A 405 means you're posting to an SSE server's /sse. A 400 on connect from an SSE client usually means the URL is a Streamable HTTP /mcp. And "Session terminated" on the very first connect usually means the path is wrong: a Streamable HTTP URL ends in /mcp.


[05]

Run an MCP server over HTTP in production

Streamable HTTP is the transport to deploy. These steps take the inventory server from your laptop to something you'd put behind a domain name.

Step 01

Keep the server on loopback, and know what that protects

The MCP specification says servers must validate the Origin header, because of DNS rebinding: a web page in your browser can re-point its own hostname at 127.0.0.1 and talk to a server you thought was private. Promptise checks Host and Origin before any MCP message is parsed. On a loopback bind, the checks are on by default. These are steps 5 and 6 of http_raw.sh, against the server from Step 3:

>_Terminalhttp_raw.sh
echo; echo; echo "### 5. a foreign Host header (DNS rebinding)"
curl -si $URL -H "$ACCEPT" -H "$JSON" -H "Host: attacker.example" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'

echo; echo; echo "### 6. a foreign Origin header"
curl -si $URL -H "$ACCEPT" -H "$JSON" -H "Origin: http://attacker.example" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'
Output
### 5. a foreign Host header (DNS rebinding)
HTTP/1.1 421 Misdirected Request
…
Invalid Host header

### 6. a foreign Origin header
HTTP/1.1 403 Forbidden
…
Invalid Origin header

Both refused before a session was created. Requests without an Origin header, which is every non-browser MCP client, pass the Origin check.

Warning

server.run() defaults to host="0.0.0.0", and on a non-loopback bind Promptise turns the Host check off unless you pass allowed_hosts. The same server bound to 0.0.0.0 with no list answered Host: attacker.example with 200. Always pass host explicitly. The promptise serve command defaults to 127.0.0.1 instead.

Step 02

Configure the server from the environment

A production server needs three more things: the public hostname your proxy will forward, an API key on every request, and a readiness check. Keep all of them out of the code:

Pythonprod_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
import json
import logging
import os
from pathlib import Path

from promptise.mcp.server import APIKeyAuth, AuthMiddleware, HealthCheck, MCPServer

# INFO shows the Host/Origin policy at startup, so you can check a deployment from its logs.
logging.basicConfig(level=logging.INFO)

STOCK_FILE = Path(os.environ.get("STOCK_FILE", "stock.json"))

server = MCPServer("inventory", version="1.0.0", require_auth=True)
server.add_middleware(AuthMiddleware(APIKeyAuth(keys={os.environ["INVENTORY_API_KEY"]: "agents"})))


@server.tool(read_only_hint=True)
async def check_stock(sku: str) -> dict:
    """Get how many units of a product are in stock.

    Args:
        sku: The product code, for example "SKU-1".
    """
    stock = json.loads(STOCK_FILE.read_text())
    sku = sku.strip().upper()
    if sku not in stock:
        return {"error": f"Unknown product {sku}."}
    return {"sku": sku, "in_stock": stock[sku]}


# Liveness and readiness, served as MCP resources health://liveness and health://readiness.
health = HealthCheck()
health.add_check("stock_file", STOCK_FILE.is_file, required_for_ready=True)
health.register_resources(server)

if __name__ == "__main__":
    server.run(
        transport="http",
        host=os.environ.get("HOST", "127.0.0.1"),
        port=int(os.environ.get("PORT", "8080")),
        # The public name your proxy forwards as Host, e.g. "mcp.example.com".
        allowed_hosts=os.environ["ALLOWED_HOSTS"].split(","),
    )

With require_auth=True and an AuthMiddleware, Promptise rejects any HTTP request without a valid x-api-key or bearer token before it reaches MCP. Start it with HOST=127.0.0.1 PORT=8174 ALLOWED_HOSTS=mcp.example.com python prod_server.py, and right after the startup banner the log tells you the policy it's enforcing:

Output
INFO:promptise.server:Host/Origin validation on: hosts=['127.0.0.1', '127.0.0.1:*', 'localhost', 'localhost:*', '[::1]', '[::1]:*', 'mcp.example.com'] origins=[…]
INFO:promptise.server:MCP server running on http://127.0.0.1:8174/mcp

Without logging.basicConfig(level=logging.INFO) that line isn't printed at all, so keep it. Authentication & Security covers JWTs and roles if API keys aren't enough.

Step 03

Put it behind a reverse proxy with TLS

Don't terminate TLS in the MCP server. Put Nginx, Caddy or your cloud's load balancer in front, let it handle certificates, and have it forward the public Host header to the server. What your server accepts then depends on where it listens, and the two common setups behave differently. This script sends the Host header each proxy would forward:

>_Terminalprod_hosts.sh
# Which Host headers does each bind accept? INVENTORY_API_KEY is set in the environment.
#   8174: HOST=127.0.0.1 ALLOWED_HOSTS=mcp.example.com python prod_server.py   (behind a proxy on the same machine)
#   8175: HOST=0.0.0.0   ALLOWED_HOSTS=mcp.example.com python prod_server.py   (in a container)
#   8176: HOST=0.0.0.0   python inventory_server.py http 8176                  (public bind, no allowed_hosts)
INIT='{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'

probe() {  # probe <port> <Host header> [extra curl args]
  port=$1; host=$2; shift 2
  code=$(curl -s -o /dev/null -w "%{http_code}" http://127.0.0.1:$port/mcp \
    -H "Accept: application/json, text/event-stream" -H "Content-Type: application/json" \
    -H "Host: $host" "$@" -d "$INIT")
  printf "%-5s Host: %-22s -> %s\n" "$port" "$host" "$code"
}

KEY=(-H "x-api-key: $INVENTORY_API_KEY")
probe 8174 mcp.example.com "${KEY[@]}"
probe 8174 127.0.0.1:8174 "${KEY[@]}"
probe 8174 attacker.example "${KEY[@]}"
probe 8174 mcp.example.com
echo
probe 8175 mcp.example.com "${KEY[@]}"
probe 8175 127.0.0.1:8175 "${KEY[@]}"
probe 8175 attacker.example "${KEY[@]}"
echo
probe 8176 attacker.example
Output
8174  Host: mcp.example.com        -> 200
8174  Host: 127.0.0.1:8174         -> 200
8174  Host: attacker.example       -> 421
8174  Host: mcp.example.com        -> 401

8175  Host: mcp.example.com        -> 200
8175  Host: 127.0.0.1:8175         -> 421
8175  Host: attacker.example       -> 421

8176  Host: attacker.example       -> 200

Read it row by row:

  • Proxy on the same machine (8174). Bind 127.0.0.1 and list the public name. allowed_hosts adds to the loopback names, so the proxied traffic and your local checks both work. The wrong host gets 421, a missing key gets 401.

  • In a container (8175). You have to bind 0.0.0.0 so traffic can reach the container. On that bind, allowed_hosts is used exactly as given, so even 127.0.0.1 is refused. Add 127.0.0.1:* if anything inside the container, like a health check, calls the server locally.

  • No list at all (8176). Every Host is accepted. Only do this if a gateway in front validates Host and Origin for you.

Whatever proxy you use, check four settings: forward the original Host header, pass mcp-session-id through untouched, don't buffer text/event-stream responses, and allow long read timeouts for slow tools. The Deployment docs have an Nginx example. It wasn't run for this guide.

Step 04

Add a health check your orchestrator can run

Kubernetes, Docker and load balancers want to ask "are you ready?" The obvious probe, a plain GET /mcp, doesn't work. It's the last step of http_raw.sh:

Output
### 9. a plain GET, like a naive health probe
HTTP/1.1 406 Not Acceptable
…
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Not Acceptable: Client must accept text/event-stream"}}

A probe that treats 4xx as a failure would restart a healthy server forever. Promptise's HealthCheck serves liveness and readiness as MCP resources rather than plain HTTP routes, so the probe needs to speak MCP. This script does, and turns the answer into an exit code:

Pythonhealthcheck.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
"""Exit 0 when the MCP server is ready, 1 otherwise. Use it as a container or exec probe."""

import asyncio
import json
import os
import sys

from promptise.mcp.client import MCPClient

URL = f"http://127.0.0.1:{os.environ.get('PORT', '8080')}/mcp"


async def main() -> int:
    try:
        async with MCPClient(url=URL, api_key=os.environ["INVENTORY_API_KEY"], timeout=5) as client:
            result = await client.session.read_resource("health://readiness")
    except Exception as exc:
        print(f"unhealthy: {exc}")
        return 1
    report = json.loads(result.contents[0].text)
    print(json.dumps(report))
    return 0 if report["status"] == "ready" else 1


sys.exit(asyncio.run(main()))

Here it is against the server on 8174, then with the stock file moved away, then against the container-style server on 8175:

Output
$ PORT=8174 python healthcheck.py; echo "exit $?"
{"status": "ready", "checks": {"stock_file": {"healthy": true}}}
exit 0
$ mv stock.json stock.json.bak
$ PORT=8174 python healthcheck.py; echo "exit $?"
{"status": "not_ready", "checks": {"stock_file": {"healthy": false}}}
exit 1
$ PORT=8175 python healthcheck.py; echo "exit $?"
unhealthy: Server at http://127.0.0.1:8175/mcp rejected the connection: 421 Misdirected Request.
exit 1

The last line is the container trap from Step 3. After restarting that server with ALLOWED_HOSTS="mcp.example.com,127.0.0.1:*", the same probe passes:

Output
$ PORT=8175 python healthcheck.py; echo "exit $?"
{"status": "ready", "checks": {"stock_file": {"healthy": true}}}
exit 0

Use the script as an exec probe in Kubernetes or a HEALTHCHECK CMD in Docker. For pure liveness, a TCP check on the port is enough. Docker wasn't available on the machine used for this guide, so no container image was built; python prod_server.py and python healthcheck.py are the two commands a container would run.

Step 05

Plan for restarts and more than one replica

Sessions live in the memory of one server process. Two consequences follow, and both are worth seeing before your first deploy. First, run two copies of the server and open a session on one of them:

>_Terminaltwo_replicas.sh
# Open a session on replica A.
SESSION=$(curl -si http://127.0.0.1:8172/mcp -H "$ACCEPT" -H "$JSON" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}' \
  | grep -i '^mcp-session-id:' | cut -d' ' -f2 | tr -d '\r')
echo "session from replica A: $SESSION"
curl -s -o /dev/null http://127.0.0.1:8172/mcp -H "$ACCEPT" -H "$JSON" -H "mcp-session-id: $SESSION" \
  -d '{"jsonrpc":"2.0","method":"notifications/initialized"}'

for replica in A:8172 B:8173; do
  echo "--- tools/list on replica ${replica%%:*}"
  curl -s -w "\nHTTP %{http_code}\n" http://127.0.0.1:${replica##*:}/mcp -H "$ACCEPT" -H "$JSON" \
    -H "mcp-session-id: $SESSION" -d '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' | cut -c1-120
done
Output
session from replica A: 417143059734432793edf38c94c06910
--- tools/list on replica A
event: message
data: {"jsonrpc":"2.0","id":2,"result":{"tools":[{"name":"check_stock","description":"Get how many units of a product ar


HTTP 200
--- tools/list on replica B
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Session not found"}}
HTTP 404

Replica B has never heard of the session. Behind a round-robin load balancer, half of a client's requests would fail like this. Turn on sticky sessions (session affinity) so every request from one client reaches the same replica, or run a single replica per server until you need more.

Second, restart the server while an agent is connected, which is what every deploy does. This script asks the same question before and after a restart:

Pythonagent_restart.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
async def main():
    server = start_server()
    await wait_until_up()
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"inventory": HTTPServerSpec(url="http://127.0.0.1:8172/mcp")},
        instructions="You answer stock questions. Use your tools. Reply in one sentence.",
        trace_tools=True,
    )
    try:
        await ask(agent, "before restart")
        server.terminate(); server.wait()
        server = start_server()
        await wait_until_up()
        await ask(agent, "after restart")
    finally:
        await agent.shutdown()
        server.terminate(); server.wait()
Output
→ Invoking tool: check_stock with {'sku': 'SKU-1'}
✔ Tool result from check_stock: {"sku": "SKU-1", "in_stock": 42}
[before restart] There are 42 units of SKU-1 in stock.
→ Invoking tool: check_stock with {'sku': 'SKU-1'}
✖ check_stock error: Failed to call tool 'check_stock': Session terminated
→ Invoking tool: check_stock with {'sku': 'SKU-1'}
✖ check_stock error: Unknown tool 'check_stock'. Call list_tools() first to discover tools.
…
[after restart] I couldn't retrieve the stock level for SKU-1 due to an internal tool error; please try again or tell me if you'd like me to check a different SKU.

The new server process has no sessions, so the agent's session is gone. In Promptise 1.2.1 the agent doesn't open a new one by itself: every later call fails until you build the agent again. A fresh connection works immediately. For a long-lived service, catch this and rebuild the agent with build_agent, or create agents per job rather than once at startup.


[06]

Browser clients and CORS

If a web app on another origin calls your server directly, the browser needs CORS headers, and Promptise's Origin check needs to allow that origin too. Configure both:

Pythoncors_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
from promptise.mcp.server import CORSConfig

from inventory_server import server

server.run(
    transport="http",
    host="127.0.0.1",
    port=8178,
    allowed_origins=["https://app.example.com"],
    cors=CORSConfig(
        allow_origins=["https://app.example.com"],
        allow_headers=["Content-Type", "Accept", "mcp-session-id", "mcp-protocol-version"],
    ),
)

List mcp-session-id and mcp-protocol-version in allow_headers yourself. The defaults are Content-Type, Authorization and x-api-key, and with them the browser's preflight for a request that carries a session ID gets 400 Disallowed CORS headers.

There's one gap you can't configure away in 1.2.1. A browser script can only read the mcp-session-id response header if the server sends Access-Control-Expose-Headers: mcp-session-id, and CORSConfig has no option for it. Here's an initialize sent from that origin to the server above:

Output
### initialize from the browser origin
HTTP/1.1 200 OK
…
content-type: text/event-stream
mcp-session-id: 5c9dc06defeb48878fe1b2a64a7728e8
x-accel-buffering: no
access-control-allow-origin: https://app.example.com
vary: Origin
transfer-encoding: chunked

The session ID is there, but without an expose header the browser hides it from your code, so the session can't continue. Serve the web app and the MCP endpoint from the same origin through your proxy, or have the proxy add that header.


[07]

Honest limits

  • Sessions are in memory only. Promptise 1.2.1 runs Streamable HTTP with sessions on and no event store, and doesn't offer a stateless mode. Scale out with sticky sessions, and expect every restart to end every session.

  • Agents don't reconnect after a server restart. Rebuild the agent, as Step 5 shows. The MCP spec expects a client to start a new session after a 404.

  • Health checks are MCP resources, not HTTP routes. There's no /health URL; probe with a small MCP client like healthcheck.py, or a TCP check for liveness.

  • One process per `server.run()`. Promptise starts its own uvicorn server and doesn't hand you the ASGI app, so you can't add uvicorn workers. Run more replicas instead.

  • Browser clients across origins can't read the session ID without a proxy-added expose header.

  • Not tested here: a real Nginx or Caddy in front, and a Docker build. The proxy behaviour was reproduced by sending the Host header a proxy forwards.


[08]

Frequently asked questions

What is the difference between stdio and Streamable HTTP in MCP?

With stdio, the client starts your server as a child process and they exchange JSON-RPC messages over its standard input and output, one per line. With Streamable HTTP, your server runs on its own and clients send each message as an HTTP POST to one /mcp endpoint. stdio suits a local tool for one user; Streamable HTTP suits anything shared, remote or behind authentication.

Is SSE deprecated in MCP?

Yes. MCP protocol version 2025-03-26 replaced the HTTP with SSE transport from 2024-11-05 with Streamable HTTP, and clients like Claude Code now label SSE deprecated. Promptise still serves and connects over SSE with transport="sse", so you can support older clients, but build new servers on Streamable HTTP.

What does "Missing session ID" mean on an MCP server?

Your client sent a request to a Streamable HTTP server without the mcp-session-id header it received from initialize, and the server answered 400 Bad Request: Missing session ID. Either the client isn't sending the header back, or a proxy is dropping it. A 404 Session not found is different: the session existed, but the server restarted, ended it, or the request landed on another replica.

Is "http" the same as "streamable-http" in Promptise?

Yes. In Promptise Foundry 1.2.1, HTTPServerSpec(transport="http") and transport="streamable-http" both use the MCP SDK's Streamable HTTP client, and "http" is the default. Use transport="sse" only for legacy SSE servers, with a URL ending in /sse.

How do I deploy an MCP server to production?

Run it with transport="http" and an explicit host, put a reverse proxy or load balancer in front for TLS, and pass the public hostname in allowed_hosts. Require an API key or bearer token, probe readiness with an MCP client, and use sticky sessions if you run more than one replica. Build a Production MCP Server in Python covers what goes inside the server: validation, roles, rate limits and audit logs.


[09]

Where to go next

  • Deployment: transports, Host and Origin validation, CORS, proxies and the promptise serve command.

  • Server fundamentals: tools, resources and prompts.

  • Server configuration: every option on StdioServerSpec and HTTPServerSpec.

  • MCP client: MCPClient and MCPMultiClient without an agent.

  • Resilience patterns: HealthCheck, circuit breakers and more.

  • MCP transports specification: the protocol revision Promptise 1.2.1 speaks.

  • How to Connect MCP Servers to Your AI Agent in Python and OpenAPI to MCP: Turn Any REST API into an MCP Server.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths