PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Streaming LLM Responses with FastAPI: Tokens and Tool Calls

Stream an AI agent's tool calls and answer from FastAPI to a web page with Server-Sent Events, with cancellation, errors and an EventSource page.

Level
Intermediate
Reading time
14 min
Published
Oct 10, 2026
By
Promptise Team
  • Streaming
  • FastAPI
  • Server-Sent Events
  • AI Agents
  • Tool Calling
  • Promptise Foundry

An agent that calls tools can take ten seconds or more to answer, and a blank screen for ten seconds feels broken. Streaming fixes that: the page shows "Getting order status…" the moment the agent calls a tool, then the result, then the answer. In this guide you'll stream an AI agent's response from FastAPI to a web page with Server-Sent Events: tool-call events, tokens, cancellation when the user leaves, and errors halfway through. Every snippet ran against Promptise Foundry 1.2.1, and every output is pasted from a real run, including the places where 1.2.1 doesn't stream yet.

[01]

How do you stream an LLM response with FastAPI?

Return a StreamingResponse with the media type text/event-stream, and write each event from your agent as a data: line followed by a blank line. That format is Server-Sent Events (SSE), and every browser can read it with EventSource. With Promptise Foundry, agent.astream_with_tools() gives you the events: one when a tool starts, one when it ends, one per token, and one when the agent is done.

Pythonapp.py
1
2
3
4
5
6
7
8
9
10
11
12
13
@app.get("/chat")
async def chat(q: str):
    async def sse():
        async for event in app.state.agent.astream_with_tools(
            {"messages": [{"role": "user", "content": q}]}
        ):
            yield f"data: {event.to_json()}\n\n"

    return StreamingResponse(
        sse(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )

Tool events stream live, and so do tokens when the agent answers without a tool. After a tool call, Promptise 1.2.1 doesn't send the answer at all. Step 4 adds a small fix that delivers it with the final event.


[02]

How agent streaming works

The browser opens one long HTTP response. Your endpoint loops over the agent's events and writes each one as soon as it exists. Nothing waits for the whole run to finish. When the model answers without calling a tool, token events stream between the start and done.

Rendering diagram…

These are the methods on the agent, printed from the installed package:

Output
astream (self, input: 'Any', config: 'dict[str, Any] | None' = None, *, caller: 'CallerContext | None' = None, **kwargs: 'Any') -> 'AsyncIterator[Any]'
astream_with_tools (self, input: 'Any', config: 'dict[str, Any] | None' = None, *, caller: 'CallerContext | None' = None, include_arguments: 'bool' = True, tool_display_names: 'dict[str, str] | None' = None, **kwargs: 'Any') -> 'AsyncIterator[Any]'

Use astream_with_tools(). In 1.2.1, astream() fails before it yields anything, with AttributeError: 'PromptGraphEngine' object has no attribute 'astream'.

astream_with_tools() yields event objects. All of them share type and timestamp from the StreamEvent base class, and to_json() turns any of them into one SSE-ready line:

type

Class

When

Fields

tool_start

ToolStartEvent

The agent calls a tool

tool_name, tool_display_name, arguments, tool_index

tool_end

ToolEndEvent

The tool returns

tool_name, tool_summary, duration_ms, success, tool_index

token

TokenEvent

The model writes text

text, cumulative_text

done

DoneEvent

The run finished

full_response, tool_calls, duration_ms, cache_hit

error

ErrorEvent

The run failed

message, recoverable

timestamp comes from the server's monotonic clock, so it's only useful for measuring the time between two events.


[03]

What you need

  • Python 3.10 or newer. This guide ran on Python 3.12 with FastAPI 0.143.0 and uvicorn 0.54.0.

  • Promptise Foundry, FastAPI and uvicorn: pip install "promptise==1.2.1" fastapi "uvicorn[standard]"

  • An OpenAI API key in OPENAI_API_KEY. Other providers work by changing the model string, as listed in Models & Providers.

  • curl, to watch the raw stream.


[04]

Stream an agent to a web page, step by step

You'll build a support assistant with one MCP tool that looks up orders, put it behind a FastAPI endpoint, and show its progress on a page.

Step 01

Write an MCP server with one slow tool

Real tools take time, and that time is what the user stares at. This tool sleeps for two seconds to stand in for a slow warehouse API, and prints to stderr when it starts and finishes so you can see it working from the server log.

Pythonorders_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
import asyncio
import sys

from promptise.mcp.server import MCPServer, ToolError

server = MCPServer("orders")

# A stand-in for your real database or API.
ORDERS = {
    "A-1001": {"status": "shipped", "carrier": "DHL", "eta": "2026-10-14"},
    "A-1002": {"status": "processing", "carrier": None, "eta": None},
}


@server.tool()
async def get_order_status(order_id: str) -> dict:
    """Get an order's status, carrier and expected delivery date.

    Args:
        order_id: The order ID, for example "A-1001".
    """
    print(f"[orders] lookup {order_id} started", file=sys.stderr, flush=True)
    await asyncio.sleep(2)  # pretend the warehouse API is slow
    order = ORDERS.get(order_id.strip().upper())
    if order is None:
        raise ToolError(f"No order found with ID {order_id}.")
    print(f"[orders] lookup {order_id} finished", file=sys.stderr, flush=True)
    return {"order_id": order_id.strip().upper(), **order}


if __name__ == "__main__":
    server.run()

If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python explains each part of this file.

Step 02

Write the FastAPI endpoint

Build the agent once when the app starts, not once per request: connecting to MCP servers takes time, and one agent can serve many requests at once. Two streams I ran side by side each got their own tool events and their own answer. FastAPI's lifespan is the place for that.

Pythonapp.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
import sys
from contextlib import asynccontextmanager
from pathlib import Path

from fastapi import FastAPI
from fastapi.responses import StreamingResponse

from promptise import build_agent
from promptise.config import StdioServerSpec

HERE = Path(__file__).parent


@asynccontextmanager
async def lifespan(app: FastAPI):
    # Build the agent once at startup and share it between requests.
    app.state.agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "orders": StdioServerSpec(command=sys.executable, args=[str(HERE / "orders_server.py")]),
        },
        instructions="You are a friendly support assistant. Use your tools to answer questions about orders.",
    )
    yield
    await app.state.agent.shutdown()


app = FastAPI(lifespan=lifespan)


@app.get("/chat")
async def chat(q: str):
    async def sse():
        async for event in app.state.agent.astream_with_tools(
            {"messages": [{"role": "user", "content": q}]}
        ):
            yield f"data: {event.to_json()}\n\n"

    return StreamingResponse(
        sse(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )

Three choices here are deliberate:

  • The endpoint is a `GET`. The browser's EventSource can only send GET requests, so the question travels in the query string.

  • Each event ends with a blank line. That's how SSE separates events. Forget the second \n and the browser never fires a message.

  • Two headers stop caching and buffering. Cache-Control: no-cache keeps caches out of the way, and X-Accel-Buffering: no tells nginx not to hold the response back. More on that under Before you go to production.

Step 03

Run it and watch the stream with curl

Start the app, then call it from a second terminal. -N tells curl to print data as it arrives instead of buffering it:

>_Terminal
uvicorn app:app --port 8220
curl -sN "http://127.0.0.1:8220/chat?q=Where%20is%20my%20order%20A-1001%3F"
Output
data: {"type": "tool_start", "timestamp": 30856.453926208, "tool_name": "get_order_status", "tool_display_name": "Getting order status", "arguments": {"order_id": "A-1001"}, "tool_index": 0}

data: {"type": "tool_end", "timestamp": 30858.748135333, "tool_name": "get_order_status", "tool_summary": "order_id: A-1001, status: shipped, carrier: DHL (+1 more)", "duration_ms": 5095.1, "success": true, "tool_index": 0}

data: {"type": "done", "timestamp": 30863.280884833, "full_response": "", "tool_calls": [{"name": "get_order_status", "summary": "order_id: A-1001, status: shipped, carrier: DHL (+1 more)", "success": true}], "duration_ms": 9628.2, "cache_hit": false}

The tool events arrive the moment they happen, each with a ready-made label (Getting order status) and a one-line summary of the result. That part is exactly what a progress UI needs.

Now look at full_response. It's empty, and no token events came at all. In Promptise 1.2.1, the model call that writes the answer after a tool call runs without streaming, and its text never reaches the event stream. The model did write an answer: a LangChain callback on the same run saw it come back. The stream just doesn't carry it.

Ask something that needs no tool, and tokens do stream:

>_Terminal
curl -sN "http://127.0.0.1:8220/chat?q=Hi%21%20In%20one%20sentence%2C%20what%20can%20you%20help%20me%20with%3F"
Output
data: {"type": "token", "timestamp": 30873.779240958, "text": "I", "cumulative_text": "I"}

data: {"type": "token", "timestamp": 30873.78222325, "text": " can", "cumulative_text": "I can"}

data: {"type": "token", "timestamp": 30873.783247916, "text": " help", "cumulative_text": "I can help"}
…
data: {"type": "token", "timestamp": 30873.956835291, "text": ".", "cumulative_text": "I can help with support for your orders \u2014 checking order status and tracking, finding delivery dates, and answering other order-related questions."}

data: {"type": "done", "timestamp": 30877.915742208, "full_response": "I can help with support for your orders \u2014 checking order status and tracking, finding delivery dates, and answering other order-related questions.", "tool_calls": [], "duration_ms": 5610.3, "cache_hit": false}

text is the new piece and cumulative_text is everything so far, so a client that misses an event can still redraw the full text. Notice the four seconds between the last token and done, too. Honest limits explains where they go.

Step 04

Make sure the answer reaches the page

Until the answer after a tool call streams, catch it yourself. astream_with_tools() passes its config through to LangChain, so a small callback handler can remember the text of the last model call. When done arrives with an empty full_response, fill it in before you send it. These are the lines that change in app.py:

Pythonapp.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
from langchain_core.callbacks import AsyncCallbackHandler


class LastModelOutput(AsyncCallbackHandler):
    """Remembers the text of the most recent model call."""

    def __init__(self):
        self.text = ""

    async def on_llm_end(self, response, **kwargs):
        self.text = response.generations[0][0].message.content


@app.get("/chat")
async def chat(q: str):
    async def sse():
        last = LastModelOutput()
        try:
            async for event in app.state.agent.astream_with_tools(
                {"messages": [{"role": "user", "content": q}]},
                config={"callbacks": [last]},
            ):
                # Promptise 1.2.1: after a tool call the answer isn't streamed,
                # so fill it in from the last model call.
                if event.type == "done" and not event.full_response:
                    event.full_response = last.text
                yield f"data: {event.to_json()}\n\n"
        except asyncio.CancelledError:
            print(f"client disconnected, agent run stopped: {q!r}", flush=True)
            raise

The except asyncio.CancelledError block is for Step 6; add import asyncio at the top of the file. Run the same question again, this time with -i to see the headers:

>_Terminal
curl -isN "http://127.0.0.1:8220/chat?q=Where%20is%20my%20order%20A-1001%3F"
Output
HTTP/1.1 200 OK
date: Sat, 10 Oct 2026 16:47:05 GMT
server: uvicorn
cache-control: no-cache
x-accel-buffering: no
content-type: text/event-stream; charset=utf-8
transfer-encoding: chunked

data: {"type": "tool_start", "timestamp": 30983.319831666, "tool_name": "get_order_status", "tool_display_name": "Getting order status", "arguments": {"order_id": "A-1001"}, "tool_index": 0}

data: {"type": "tool_end", "timestamp": 30985.583999833, "tool_name": "get_order_status", "tool_summary": "order_id: A-1001, status: shipped, carrier: DHL (+1 more)", "duration_ms": 3800.5, "success": true, "tool_index": 0}

data: {"type": "done", "timestamp": 30989.239597916, "full_response": "Your order A-1001 has shipped via DHL and is currently in transit. The expected delivery date is 2026-10-14.…", …}

The answer is there now. It arrives in one piece with done instead of word by word, which is the honest best 1.2.1 can do after a tool call. The page in the next step handles both cases: it appends tokens as they come, and replaces the text with full_response at the end.

Note

Only fill in full_response when it's empty. When the agent answers without a tool, the streamed tokens are the answer, and the last model call holds a second, discarded version of it.

Step 05

Add a web page that reads the stream

The browser side is plain JavaScript. EventSource opens the stream, and one handler sorts events by type: tool events become a list of steps, tokens are appended to the answer, and done or error ends it.

Textindex.html
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <title>Order assistant</title>
  <style>
    body { font-family: system-ui, sans-serif; max-width: 40rem; margin: 2rem auto; padding: 0 1rem; }
    #steps { color: #555; font-size: 0.9rem; padding-left: 1.2rem; }
    #answer { white-space: pre-wrap; line-height: 1.5; }
  </style>
</head>
<body>
  <form id="ask">
    <input id="q" size="40" value="Where is my order A-1001?">
    <button>Ask</button>
  </form>
  <ul id="steps"></ul>
  <div id="answer"></div>

  <script>
    const steps = document.getElementById("steps");
    const answer = document.getElementById("answer");
    let source = null;

    document.getElementById("ask").addEventListener("submit", (e) => {
      e.preventDefault();
      if (source) source.close();
      steps.innerHTML = "";
      answer.textContent = "";

      const q = document.getElementById("q").value;
      source = new EventSource("/chat?q=" + encodeURIComponent(q));

      source.onmessage = (msg) => {
        const event = JSON.parse(msg.data);
        if (event.type === "tool_start") {
          const li = document.createElement("li");
          li.textContent = event.tool_display_name + "…";
          steps.appendChild(li);
        } else if (event.type === "tool_end") {
          const li = document.createElement("li");
          li.textContent = "Done: " + event.tool_summary;
          steps.appendChild(li);
        } else if (event.type === "token") {
          answer.textContent += event.text;
        } else if (event.type === "done") {
          answer.textContent = event.full_response;
          source.close();
        } else if (event.type === "error") {
          answer.textContent = event.message;
          source.close();
        }
      };

      // Network failure: stop instead of letting the browser retry the question.
      source.onerror = () => source.close();
    });
  </script>
</body>
</html>

Serve it from the same app, so the page and the stream share an origin and you need no CORS setup. Add this route to app.py, and HTMLResponse to the fastapi.responses import:

Pythonapp.py
@app.get("/")
async def index():
    return HTMLResponse((HERE / "index.html").read_text())
>_Terminal
curl -is http://127.0.0.1:8220/
Output
HTTP/1.1 200 OK
date: Sat, 10 Oct 2026 16:47:16 GMT
server: uvicorn
content-length: 1950
content-type: text/html; charset=utf-8

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <title>Order assistant</title>
…

Open http://127.0.0.1:8220/ in your browser and ask a question.

The source.close() calls matter more than they look. When a stream ends, EventSource reconnects on its own after a few seconds, and here a reconnect means the same question goes to the agent again. I checked with Node's built-in EventSource, which follows the same standard, connected to this endpoint and never closed:

Output
0.1s connection opened (#1)
1.5s tool_start
3.6s tool_end
4.9s done
5.0s stream ended, readyState = 0
8.1s connection opened (#2)
8.1s reconnected and asked the same question again; closing

Three seconds after done, it was back and the agent started a second run. Close the source on done and on error, every time.

Step 06

Stop the agent when the user leaves

When a user closes the tab, there's no point paying for the rest of the run. With uvicorn, Starlette notices the disconnect and cancels your generator, which raises asyncio.CancelledError inside the async for loop. Step 4's except block logs it and re-raises, so the cancellation carries on through the agent.

To test it, disconnect curl after three seconds, while the tool is still running:

>_Terminal
curl -sN --max-time 3 "http://127.0.0.1:8220/chat?q=Where%20is%20my%20order%20A-1002%3F"

The server log:

Output
INFO:     127.0.0.1:62668 - "GET /chat?q=Where%20is%20my%20order%20A-1002%3F HTTP/1.1" 200 OK
[orders] lookup A-1002 started
client disconnected, agent run stopped: 'Where is my order A-1002?'
[orders] lookup A-1002 finished

Two things to read from this. The agent run stopped: I also cancelled a run in a script right after tool_start and watched the model callbacks for ten more seconds, and no further model call started. But the tool on the MCP server still ran to the end. Cancelling the agent doesn't stop work that's already running on the server, so tools that change data should finish safely even when nobody is waiting for them. The next request on the same agent worked normally.

Tip

Always re-raise CancelledError. The Python docs warn that task groups and timeouts rely on it and can misbehave when a coroutine swallows it.


[05]

When something fails mid-stream

Once the first event is out, the HTTP status is already 200 OK, so failures have to travel as events. Promptise has two shapes for them.

A tool that reports an error arrives as a normal tool_end, with the error in the summary. Here the order doesn't exist and the tool raised ToolError:

Output
data: {"type": "tool_end", "timestamp": 31165.276600916, "tool_name": "get_order_status", "tool_summary": "error: {'code': 'TOOL_ERROR', 'message': 'No order found with ID Z-9999.', 'retryable': False}", "duration_ms": 3889.0, "success": true, "tool_index": 0}

Note "success": true. In 1.2.1, success is only false when calling the tool fails inside the agent itself. A tool that answers with an error is a successful call, and the model goes on to tell the user the order wasn't found. If your UI marks failed steps, check the summary for error: as well.

When the run itself fails, for example because the model provider rejects the request, you get an error event and no done. To produce one, I started a second copy of the app with an invalid API key:

Output
data: {"type": "error", "timestamp": 31213.494339958, "message": "An error occurred during processing.", "recoverable": false}

The message is deliberately generic, so no internal detail reaches the browser. The detail goes to your server log instead, through the promptise.engine logger:

Output
Streaming error in node 'reason': Error code: 401 - {'error': {'message': 'Incorrect API key provided: sk-not-a*****-key. You can find your API key at https://platform.openai.com/account/api-keys.', 'type': 'invalid_request_error', 'code': 'invalid_api_key', 'param': None}, 'status': 401}

Show the generic message to the user, close the EventSource, and keep the log for yourself.


[06]

Show friendlier progress labels

The automatic labels come from the tool name: get_order_status becomes Getting order status. To say something your users understand, pass tool_display_names. To keep tool arguments out of the browser, pass include_arguments=False:

Pythondisplay_names.py
1
2
3
4
5
6
7
async for event in agent.astream_with_tools(
    {"messages": [{"role": "user", "content": "Where is my order A-1001?"}]},
    tool_display_names={"get_order_status": "Checking with the warehouse"},
    include_arguments=False,
):
    if event.type == "tool_start":
        print(event.to_json())
Output
{"type": "tool_start", "timestamp": 31287.099259583, "tool_name": "get_order_status", "tool_display_name": "Checking with the warehouse", "arguments": {}, "tool_index": 0}

Hiding arguments is worth considering for any public page. They often carry customer IDs, emails or search terms that the person looking at the screen doesn't need.


[07]

Token streaming vs tool-call latency

Streaming makes waiting visible. It doesn't make the work faster. To see where the time goes, this small client prints when each event arrives:

Pythontimed_client.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
"""Print when each server-sent event arrives, in seconds since the request started."""

import json
import sys
import time

import httpx

start = time.monotonic()
tokens, last_token = 0, 0.0
with httpx.stream("GET", "http://127.0.0.1:8220/chat", params={"q": sys.argv[1]}, timeout=None) as response:
    for line in response.iter_lines():
        if not line.startswith("data: "):
            continue
        event = json.loads(line.removeprefix("data: "))
        elapsed = time.monotonic() - start
        if event["type"] == "token":
            tokens += 1
            last_token = elapsed
            if tokens == 1:
                print(f"{elapsed:5.1f}s  first token")
            continue
        if tokens:
            print(f"{last_token:5.1f}s  last token ({tokens} tokens)")
            tokens = 0
        detail = event.get("tool_display_name") or event.get("tool_summary") or event.get("message") or ""
        print(f"{elapsed:5.1f}s  {event['type']}  {detail}".rstrip())

With a tool call:

Output
  2.0s  tool_start  Getting order status
  4.3s  tool_end  order_id: A-1001, status: shipped, carrier: DHL (+1 more)
  7.6s  done

Without one:

Output
  2.0s  first token
  2.4s  last token (27 tokens)
  5.5s  done

Read the first run as three waits. For two seconds nothing arrives: the model is deciding to call a tool, and tool_start only fires once it has finished writing the call. Then the tool's own two seconds. Then three more seconds while the model writes the answer. Streaming fills the middle wait with a visible step, and that's often enough to make the page feel alive. For the first wait, show a "Thinking…" state as soon as the request goes out; no event will come sooner.

In the second run, the answer was complete at 2.4 seconds, but done came three seconds later. If your UI waits for done before it re-enables the input box, re-enable it on the last token instead.


[08]

Honest limits

Streaming in Promptise Foundry 1.2.1 has gaps you should know before you build on it. Each one was reproduced for this guide:

  • The answer after a tool call isn't streamed. No token events, and full_response on done is empty. Step 4's callback fills it in, so the answer arrives whole at the end.

  • Answers without a tool cost two model calls. The tokens stream from the first call, then a second call runs and its result is discarded. That's the gap between the last token and done, and it's paid tokens.

  • `agent.astream()` doesn't work. It raises AttributeError: 'PromptGraphEngine' object has no attribute 'astream'. Use astream_with_tools().

  • `duration_ms` on `tool_end` counts from the start of the run, not from the tool's start. Our two-second tool reported between 3.8 and 7.9 seconds across these runs. Use the timestamp difference between tool_start and tool_end instead.

  • `tool_index` on `tool_end` is always 0. With two tool calls, the second tool_end still says 0. In the streaming path tools run one after another, so pair each tool_end with the latest tool_start.

  • Tool calls in one turn run one at a time while streaming. Asked about two orders, the agent looked them up back to back, about two seconds each. ainvoke() runs independent calls in parallel.


[09]

Before you go to production

  • Turn off proxy buffering. nginx buffers responses from upstream servers by default, which turns your stream into one late chunk. The X-Accel-Buffering: no header switches that off for this response; see proxy_buffering in the nginx docs. Other proxies and CDNs have their own settings. If curl streams fine against uvicorn but not through your proxy, buffering is the first suspect.

  • Keep the agent shared, and the request separate. Build one agent at startup, as in Step 2. Pass caller= to astream_with_tools() when your users need their own identity; Multi-user systems covers how.

  • Mind the URL. EventSource sends the question in the query string, and query strings end up in access logs: uvicorn's own log above prints every question. If questions can contain personal data, keep the stream on a GET that carries a short-lived ID, and send the question itself in a POST first.

  • Plan for idle time. A run can go quiet for several seconds while the model thinks. Check that your proxy and load balancer timeouts are longer than your slowest run.


[10]

Frequently asked questions

How do I stream an LLM response to the frontend?

Send Server-Sent Events from your backend and read them with EventSource in the browser. On the server, that's a StreamingResponse with media_type="text/event-stream" that writes data: <json> and a blank line per event. On the page, one onmessage handler appends tokens and shows tool steps, and closes the source when the run is done.

Should I use SSE or WebSockets for streaming LLM responses?

SSE, unless you need the client to send messages during the same connection. A streamed answer only flows one way, and SSE does that over plain HTTP, through ordinary proxies, with no extra library in the browser. WebSockets make sense for things like voice or live collaboration, where both sides talk at once.

Can I stream with a POST request?

Not with EventSource, which only sends GET requests. The same endpoint works as a POST if you read the response with fetch() and its body stream instead, and parse the data: lines yourself. You give up the browser's built-in parsing and get a request body in return.

Why does my stream arrive all at once?

Something between the server and the client is buffering it. Test the endpoint directly with curl -N first. If that streams, look at the proxy in front: set X-Accel-Buffering: no for nginx, and check your CDN or load balancer for a buffering option.

Is Promptise a LangChain streaming alternative?

Promptise Foundry uses LangChain's model integrations underneath, so LangChain callbacks still work, as Step 4 shows. What it adds for streaming is a small set of typed events, with tool display names, result summaries and argument hiding built in, instead of raw callback events you have to filter yourself.


[11]

Where to go next

  • Streaming with Tool Visibility: the reference for astream_with_tools() and its events.

  • Building Agents: every option on build_agent.

  • Guardrails: checks on input and output, including tool arguments in the stream.

  • Observability and Events: traces and notifications to run alongside the stream.

  • Build a Production MCP Server in Python: make the tools behind your agent safe to call.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and MCP basics this guide builds on.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths