PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

MCP Long-Running Tasks: Background Jobs in Python

Run slow MCP tools as background jobs in Python: submit, poll progress, retry failures and cancel, with a real agent, real runs and honest limits.

Level
Intermediate
Reading time
19 min
Published
Oct 10, 2026
By
Promptise Team
  • MCP
  • MCP Server
  • Background Jobs
  • Python
  • Promptise Foundry

Some tools are slow. A quarterly report, a full export or a site crawl can take minutes, and an MCP tool call that takes minutes holds up the agent, the user waiting on it, and sooner or later a client timeout. The fix is the same one web backends use: start the work in the background, hand back a job ID straight away, and let the caller check on it. In this guide you'll build an MCP server in Python whose slow report runs as a background job with progress, retries and cancellation, then put a real agent in front of it. You'll also see what progress notifications, cancellation and fire-and-forget tasks do on a normal tool call, and where the limits are. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed.

[01]

How do you handle long-running tasks in an MCP server?

Don't do the slow work inside the tool call. Register it as a job on a queue, so the agent submits it, gets a job ID back in milliseconds, and checks its status or fetches its result whenever it likes. With Promptise Foundry, that's MCPQueue and the @queue.job() decorator:

Pythonquickstart_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import asyncio

from promptise.mcp.server import MCPQueue, MCPServer

server = MCPServer("reports")
queue = MCPQueue(server, max_workers=2)


@queue.job(name="generate_report", timeout=120)
async def generate_report(region: str, quarter: str) -> dict:
    """Build the quarterly sales report for one region."""
    await asyncio.sleep(8)  # the slow part: queries, joins, rendering
    return {"region": region, "quarter": quarter, "url": f"https://reports.example.com/{region}-{quarter}.pdf"}


if __name__ == "__main__":
    server.run(transport="http", host="127.0.0.1", port=8334)

Passing the server to MCPQueue adds five tools to it and starts the workers when the server starts. Submitting the eight-second report returns at once:

Output
['queue_submit', 'queue_status', 'queue_result', 'queue_cancel', 'queue_list']
0.00s {"job_id": "c1d57b4056bae0f0", "status": "pending", "job_type": "generate_report"}

The rest of this guide builds this out with progress, retries, cancellation and an agent that uses it.


[02]

How background jobs work in an MCP server

The tool call only puts the job on a queue. A pool of workers inside the server process picks jobs up and runs them, and every later call reads the job's record:

Rendering diagram…

These are the five tools the queue adds. Pass tool_prefix="jobs" to MCPQueue and they become jobs_submit, jobs_status and so on.

Tool

What it does

queue_submit

Queues a job by job_type with an args dict and an optional priority, and returns its job_id

queue_status

Returns the status, progress from 0.0 to 1.0, the latest progress message and the attempt count

queue_result

Returns the result once the job is finished, or the current status while it's still running

queue_cancel

Cancels a pending or running job

queue_list

Lists jobs, optionally filtered by status

A job is pending until a worker takes it, then running, and ends as completed, failed, timeout or cancelled.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry 1.2.1 from PyPI. The server and the queue need nothing else.

  • An API key for a model provider for the agent steps. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string.

>_Terminal
pip install promptise
export OPENAI_API_KEY="sk-..."

[04]

Build a report server with background jobs, step by step

The server makes quarterly sales reports, which take a while, and exports invoices to an accounting system that sometimes fails. The complete file, report_server.py, is in Step 4.

Step 01

Turn the slow function into a job

A job is an ordinary async function. Two extra parameters are filled in by the queue, based on their type annotation: a ProgressReporter that writes progress into the job's record, and a CancellationToken that tells the job when someone has cancelled it.

Pythonreport_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
@queue.job(name="generate_report", timeout=120)
async def generate_report(
    region: str,
    quarter: str,
    progress: ProgressReporter,
    cancel: CancellationToken,
) -> dict:
    """Build the quarterly sales report for one region."""
    data = SALES[region.lower()]
    for done, step in enumerate(STEPS):
        cancel.check()  # stop here if someone called queue_cancel
        await progress.report(done, total=len(STEPS), message=step)
        await asyncio.sleep(2)  # the slow part: queries, joins, rendering
    return {
        "region": region,
        "quarter": quarter,
        "net_revenue": data["revenue"] - data["refunds"],
        "orders": data["orders"],
        "url": f"https://reports.example.com/{region.lower()}-{quarter.lower()}.pdf",
    }
  • progress.report(done, total=4, message=...) stores done / total as a fraction, so status shows 0.25, not 25.

  • cancel.check() raises when the job has been cancelled. Call it between steps, at points where stopping is safe.

  • timeout=120 is the ceiling for one attempt. The default is 300 seconds, set by default_timeout on MCPQueue.

  • max_workers=2 on the queue means at most two jobs run at the same time. The rest wait as pending.

Step 02

Watch a job from a client

Before an agent gets involved, check the queue with a plain MCP client. Start the server with python report_server.py, then run this:

Pythonpoll_by_hand.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
import asyncio
import json
import time

from promptise.mcp.client import MCPClient


def show(result) -> dict:
    return json.loads(result.content[0].text)


async def main():
    async with MCPClient(url="http://127.0.0.1:8330/mcp") as client:
        print("Tools:", [t.name for t in await client.list_tools()])

        start = time.perf_counter()
        job = show(await client.call_tool(
            "queue_submit",
            {"job_type": "generate_report", "args": {"region": "EMEA", "quarter": "Q3"}},
        ))
        print(f"{time.perf_counter() - start:.2f}s submit ->", job)

        while True:
            await asyncio.sleep(1.5)
            status = show(await client.call_tool("queue_status", {"job_id": job["job_id"]}))
            print(f"{time.perf_counter() - start:.2f}s status ->", status)
            if status["status"] not in ("pending", "running"):
                break

        result = show(await client.call_tool("queue_result", {"job_id": job["job_id"]}))
        print(f"{time.perf_counter() - start:.2f}s result ->", json.dumps(result, indent=2))


asyncio.run(main())
Output
Tools: ['queue_submit', 'queue_status', 'queue_result', 'queue_cancel', 'queue_list', 'wait_for_job']
0.00s submit -> {'job_id': '7ff8b16ec4b0f528', 'status': 'pending', 'job_type': 'generate_report'}
1.51s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 1, 'progress_message': 'Loading orders'}
3.01s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}
4.52s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.5, 'attempts': 1, 'progress_message': 'Computing totals'}
…
9.04s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'completed', 'priority': 'normal', 'progress': 1.0, 'attempts': 1, 'progress_message': 'Rendering the PDF'}
9.04s result -> {
  "job_id": "7ff8b16ec4b0f528",
  …
  "result": {
    "region": "EMEA",
    "quarter": "Q3",
    "net_revenue": 402450.0,
    "orders": 1840,
    "url": "https://reports.example.com/emea-q3.pdf"
  }
}

The submit took no measurable time, the status climbed through the four steps, and the result arrived nine seconds later. wait_for_job in the tool list is a small tool you'll add in Step 4.

Step 03

Let an agent submit the job and come back later

The model only sees the five queue tools. Nothing in them says which job types exist or what arguments each one takes, so tell the agent in its instructions:

Pythonreport_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
import asyncio

from promptise import build_agent
from promptise.config import HTTPServerSpec

INSTRUCTIONS = """You are a reporting assistant. Reports are slow, so they run as background jobs.

Job types on the reports server:
- generate_report: args {"region": "EMEA" or "Americas", "quarter": "Q1" to "Q4"}

To produce a report, submit it with queue_submit and check it with queue_status.
When the status is completed, fetch it with queue_result. Don't wait for a
slow job: tell the user the job ID and the progress message, and check again
when they ask."""


async def main():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"reports": HTTPServerSpec(url="http://127.0.0.1:8330/mcp")},
        instructions=INSTRUCTIONS,
        trace_tools=True,
    )
    try:
        messages = [{"role": "user", "content": "I need the Q3 sales report for EMEA."}]
        result = await agent.ainvoke({"messages": messages})
        print(">>>", result["messages"][-1].content)

        await asyncio.sleep(10)  # the user does something else for a while

        messages = result["messages"] + [{"role": "user", "content": "Is it ready?"}]
        result = await agent.ainvoke({"messages": messages})
        print(">>>", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'EMEA', 'quarter': 'Q3'}, 'priority': 'normal'}
✔ Tool result from queue_submit: {"job_id": "21aa2cc272fed921", "status": "pending", "job_type": "generate_report"}
>>> I've submitted a background job to generate the Q3 sales report for EMEA.

Job ID: 21aa2cc272fed921
Status: pending
…
→ Invoking tool: queue_status with {'job_id': '21aa2cc272fed921'}
✔ Tool result from queue_status: {"job_id": "21aa2cc272fed921", "job_type": "generate_report", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 1, "progress_message": "Rendering the PDF"}
→ Invoking tool: queue_result with {'job_id': '21aa2cc272fed921'}
✔ Tool result from queue_result: {"job_id": "21aa2cc272fed921", …, "result": {"region": "EMEA", "quarter": "Q3", "net_revenue": 402450.0, "orders": 1840, "url": "https://reports.example.com/emea-q3.pdf"}}
>>> Yes — the report is ready.
…
You can download the report here:
https://reports.example.com/emea-q3.pdf
…

This is the point of a queue. The first turn ended as soon as the job was queued, so the user got an answer instead of a spinner. The job kept running between turns, and the second turn found it finished. The job ID lives in the conversation history, which is why passing result["messages"] into the next turn matters.

Step 04

Let the agent wait without polling

Sometimes the user does want to wait. With only the queue tools, an agent waits by calling queue_status again and again, and each check costs a model round trip. With instructions to keep checking until the job finished, and no other tool to wait with, this is how a run went:

Output
→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'EMEA', 'quarter': 'Q3'}, 'priority': 'normal'}
…
→ Invoking tool: queue_status with {'job_id': '64720c23a6330a2e'}
✔ Tool result from queue_status: {"job_id": "64720c23a6330a2e", "job_type": "generate_report", "status": "running", "priority": "normal", "progress": 0.5, "attempts": 1, "progress_message": "Computing totals"}
>>> Your EMEA Q3 report (job_id: 64720c23a6330a2e) is running.

Current status:
- Status: running
- Progress: 0%
- Message: "Loading orders"
…

Five status checks in 13 seconds, then it stopped at 50% and handed the decision back to the user. It even reported the first status it had seen, not the latest.

A better way is one tool that waits on the server side and returns when the job finishes or the time is up. It uses the queue's own status() method, so it reads the same record as queue_status. Here it is in the complete server:

Pythonreport_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
import asyncio
import time
from typing import Annotated

from pydantic import Field

from promptise.mcp.server import CancellationToken, MCPQueue, MCPServer, ProgressReporter

server = MCPServer("reports")
queue = MCPQueue(server, max_workers=2)

# A stand-in for your warehouse and your accounting system.
SALES = {
    "emea": {"orders": 1840, "revenue": 412_300.0, "refunds": 9_850.0},
    "americas": {"orders": 2310, "revenue": 598_120.0, "refunds": 14_200.0},
}
STEPS = ["Loading orders", "Matching refunds", "Computing totals", "Rendering the PDF"]
export_attempts: dict[str, int] = {}


@queue.job(name="generate_report", timeout=120)
async def generate_report(
    region: str,
    quarter: str,
    progress: ProgressReporter,
    cancel: CancellationToken,
) -> dict:
    """Build the quarterly sales report for one region."""
    data = SALES[region.lower()]
    for done, step in enumerate(STEPS):
        cancel.check()  # stop here if someone called queue_cancel
        await progress.report(done, total=len(STEPS), message=step)
        await asyncio.sleep(2)  # the slow part: queries, joins, rendering
    return {
        "region": region,
        "quarter": quarter,
        "net_revenue": data["revenue"] - data["refunds"],
        "orders": data["orders"],
        "url": f"https://reports.example.com/{region.lower()}-{quarter.lower()}.pdf",
    }


@queue.job(name="export_invoices", timeout=60, max_retries=2, backoff_base=2.0)
async def export_invoices(month: str) -> dict:
    """Send one month of invoices to the accounting system."""
    export_attempts[month] = export_attempts.get(month, 0) + 1
    await asyncio.sleep(1)
    if export_attempts[month] == 1:
        raise ConnectionError("accounting API answered 503 Service Unavailable")
    return {"month": month, "invoices": 214, "file": f"invoices-{month}.csv"}


@server.tool()
async def wait_for_job(
    job_id: str,
    seconds: Annotated[int, Field(ge=1, le=25, description="How long to wait at most.")] = 20,
) -> dict:
    """Wait for a background job to finish, then return its status.

    Use this instead of calling queue_status again and again. It returns as soon as
    the job is completed, failed, cancelled or timed out, or when the time is up.
    """
    deadline = time.monotonic() + seconds
    while True:
        status = await queue.status(job_id)
        if status["status"] not in ("pending", "running") or time.monotonic() >= deadline:
            return status
        await asyncio.sleep(0.5)


if __name__ == "__main__":
    server.run(transport="http", host="127.0.0.1", port=8330)

The 25-second cap keeps each call well inside client timeouts; if the job needs longer, the agent simply calls it again. The agent's instructions now list both job types and say: "When the user wants to wait for the result, call wait_for_job instead of polling queue_status". Asked for two things at once, it ran them side by side:

Output
→ Invoking tool: queue_submit with {'job_type': 'export_invoices', 'args': {'month': '2026-09'}, 'priority': 'normal'}
→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'Americas', 'quarter': 'Q3'}, 'priority': 'normal'}
✔ Tool result from queue_submit: {"job_id": "e1ada74daf6a3c9f", "status": "pending", "job_type": "export_invoices"}
✔ Tool result from queue_submit: {"job_id": "d74bcec543b0c071", "status": "pending", "job_type": "generate_report"}
→ Invoking tool: wait_for_job with {'job_id': 'e1ada74daf6a3c9f', 'seconds': 60}
→ Invoking tool: wait_for_job with {'job_id': 'd74bcec543b0c071', 'seconds': 60}
✔ Tool result from wait_for_job: Input validation error: 60 is greater than the maximum of 25
✔ Tool result from wait_for_job: Input validation error: 60 is greater than the maximum of 25
→ Invoking tool: wait_for_job with {'job_id': 'e1ada74daf6a3c9f', 'seconds': 25}
→ Invoking tool: wait_for_job with {'job_id': 'd74bcec543b0c071', 'seconds': 25}
✔ Tool result from wait_for_job: {"job_id": "e1ada74daf6a3c9f", "job_type": "export_invoices", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 2, "error": "Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s."}
✔ Tool result from wait_for_job: {"job_id": "d74bcec543b0c071", "job_type": "generate_report", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 1, "progress_message": "Rendering the PDF"}
→ Invoking tool: queue_result with {'job_id': 'e1ada74daf6a3c9f'}
→ Invoking tool: queue_result with {'job_id': 'd74bcec543b0c071'}
…
>>> Both background jobs have finished. Summary:
…
- Attempts: 2 (note: attempt 1 failed — "accounting API answered 503 Service Unavailable" — retried and completed)
…
(21.1s)

The model asked to wait 60 seconds, the le=25 limit refused it with a clear message, and it tried again with 25. After that, one wait per job was enough, and the export's retry happened without the agent having to do anything. The tool's description does the persuading: with the polling instructions from above and wait_for_job on the server, the agent checked once and then chose wait_for_job by itself.

The model still makes its own choices. In another run it waited for both jobs but fetched only the export's result, and never mentioned the report. If a result must not be dropped, check the job list yourself after the turn rather than relying on the model to.

Step 05

Retry jobs that fail

export_invoices fails on its first attempt with a ConnectionError, the kind of transient error worth retrying. max_retries=2 allows two more attempts, and backoff_base=2.0 sets the wait between them: backoff_base * 2 ** (attempt - 1), so 2 seconds, then 4. A small client that prints each change of status shows the whole cycle:

Output
 0.0s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'pending', 'priority': 'normal', 'progress': 0.0, 'attempts': 0}
 0.3s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 1}
 1.1s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'pending', 'priority': 'normal', 'progress': 0.0, 'attempts': 1, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
 3.3s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 2, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
 4.1s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'completed', 'priority': 'normal', 'progress': 1.0, 'attempts': 2, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
result: {'month': '2026-09', 'invoices': 214, 'file': 'invoices-2026-09.csv'}

What to notice:

  • Between attempts the job is `pending` again, with the failure in error. That error stays on the record after the retry succeeds, so read status to decide whether a job worked, not the presence of error.

  • Any exception triggers a retry, not only transient ones. A job submitted with a missing argument failed three times with needs_args() missing 1 required positional argument: 'quarter' before it gave up. Make jobs that can't succeed fail fast, or keep max_retries low.

  • Timeouts aren't retried. A job that runs past its timeout ends as timeout with Job timed out after 1s in error, on the first attempt.

  • During the backoff, the worker waits too. In 1.2.1 the worker that ran the failed attempt sleeps through the backoff before requeueing, so long backoffs tie up workers. Size max_workers with that in mind.

Step 06

Cancel a job

queue_cancel marks the job as cancelled and signals its CancellationToken. The job stops the next time it calls cancel.check():

Pythoncancel_by_hand.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
import asyncio
import json

from promptise.mcp.client import MCPClient


def show(result) -> dict:
    return json.loads(result.content[0].text)


async def main():
    async with MCPClient(url="http://127.0.0.1:8330/mcp") as client:
        job = show(await client.call_tool(
            "queue_submit", {"job_type": "generate_report", "args": {"region": "Americas", "quarter": "Q3"}}
        ))
        await asyncio.sleep(3)
        print("before:", show(await client.call_tool("queue_status", {"job_id": job["job_id"]})))
        print("cancel:", show(await client.call_tool("queue_cancel", {"job_id": job["job_id"]})))
        await asyncio.sleep(10)  # long after the job would have finished
        print("later: ", show(await client.call_tool("queue_result", {"job_id": job["job_id"]})))


asyncio.run(main())
Output
before: {'job_id': '8a81d9b4c31a23d6', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}
cancel: {'job_id': '8a81d9b4c31a23d6', 'status': 'cancelled'}
later:  {'job_id': '8a81d9b4c31a23d6', 'job_type': 'generate_report', 'status': 'cancelled', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}

The job was halfway through its second step. It stopped at the next cancel.check(), so it never reached "Computing totals", and ten seconds later it's still cancelled with no result.

Warning

Cancellation is cooperative. A job that never calls cancel.check() keeps running after queue_cancel, and in 1.2.1 its status then flips from cancelled to completed when it finishes. In a test, a job without a check showed cancelled right after the cancel and completed 2.5 seconds later, with its work done. Put a cancel.check() before every expensive or irreversible step.


[05]

Progress and cancellation on a normal tool call

Not everything needs a queue. A tool that takes ten or twenty seconds can stay a normal tool and still tell the client how it's doing. ProgressReporter and CancellationToken work there too, injected with Depends:

Pythondirect_tools_server.py
1
2
3
4
5
6
7
@server.tool()
async def crawl_site(pages: int, progress: ProgressReporter = Depends(ProgressReporter)) -> dict:
    """Crawl a number of pages and report progress as it goes."""
    for page in range(1, pages + 1):
        await asyncio.sleep(0.5)
        await progress.report(page, total=pages, message=f"Crawled page {page} of {pages}")
    return {"pages": pages}

Here, report() sends MCP progress notifications to the client while the call is still running. A client only receives them if it asked for them by sending a progress token with the call. The official MCP Python SDK does that when you pass a progress_callback:

Pythonprobe_direct_tools.py
1
2
3
4
5
6
7
8
9
    async def on_progress(progress, total, message):
        print(f"  progress {progress}/{total}: {message}")

    async with streamablehttp_client(URL) as (read, write, _):
        async with ClientSession(read, write) as session:
            await session.initialize()
            print("official client, with progress_callback:")
            r = await session.call_tool("crawl_site", {"pages": 3}, progress_callback=on_progress)
            print("  result:", r.content[0].text)
Output
official client, with progress_callback:
  progress 1.0/3.0: Crawled page 1 of 3
  progress 2.0/3.0: Crawled page 2 of 3
  progress 3.0/3.0: Crawled page 3 of 3
  result: {"pages": 3}
…
Promptise MCPClient:
  result: {"pages": 3} (no progress lines above)

Promptise's own MCP client, which a Promptise agent uses, doesn't ask for progress in 1.2.1, so the same calls send nothing to it. The model never sees progress notifications either way; they're for a person watching a client UI. If the agent needs to know how far along something is, use a queue job, whose progress the agent can read with queue_status.

Cancelling a normal tool call works differently from cancelling a job. When the client sends an MCP cancellation, the MCP SDK cancels the running handler outright. The CancellationToken isn't set. A tool that waited on its token logged this when the official client cancelled it:

Output
19:30:54 watch_feed: asyncio task cancelled, token.is_cancelled=False

The client got McpError Request cancelled. So on a normal tool, the token won't tell you about a cancellation in 1.2.1. Put any cleanup in a try/finally or catch asyncio.CancelledError, and keep cancel.check() for queue jobs, where it does work.


[06]

Fire-and-forget work with BackgroundTasks

BackgroundTasks is for small jobs after a tool has done its real work, such as sending a receipt or writing an analytics event. You add callables during the tool call and they run once the handler returns:

Pythondirect_tools_server.py
1
2
3
4
5
6
7
8
9
10
11
async def send_receipt(order_id: str) -> None:
    await asyncio.sleep(3)  # a slow email API
    print(f"{time.strftime('%X')} receipt sent for {order_id}", flush=True)


@server.tool()
async def place_order(order_id: str, bg: BackgroundTasks = Depends(BackgroundTasks)) -> dict:
    """Place an order and email the receipt afterwards."""
    bg.add(send_receipt, order_id)
    print(f"{time.strftime('%X')} handler returned for {order_id}", flush=True)
    return {"order_id": order_id, "status": "placed"}

Errors in background tasks are logged and never reach the caller. But look at the timing over HTTP. The server printed:

Output
19:30:48 handler returned for A-1001
19:30:51 receipt sent for A-1001

And the client calling place_order printed:

Output
19:30:51 place_order answered after 3.01s: {"order_id": "A-1001", "status": "placed"}

The handler returned at once, yet the caller waited three seconds for its answer. In 1.2.1, background tasks run before the response goes out, not after it, so they add their full duration to the call. The docs describe the opposite. Until that changes, use BackgroundTasks only for work that takes milliseconds, and make anything slower a queue job.

Here's how to choose:

The work takes

Use

The agent sees

A few seconds

A normal tool

The result

Up to a minute or so

A normal tool with ProgressReporter

The result; a client UI can show progress

Minutes, or it may fail and need retries

An MCPQueue job

A job ID now, status and result later

Milliseconds of follow-up work

BackgroundTasks

The result, after the tasks finish


[07]

Honest limits

These are true of Promptise Foundry 1.2.1. Each one was checked in a run or in the source.

  • Jobs live in memory. The only backend that ships is InMemoryQueueBackend. When the server restarts, every job and result is gone: asked about a job submitted before the restart, the server answered JOB_NOT_FOUND. There's no persistence and no Redis backend. You can write your own against the QueueBackend protocol, seven async methods from enqueue to count, but nothing re-runs jobs that were running when a process died.

  • One queue per process. With two replicas, a job submitted to one is JOB_NOT_FOUND on the other. Behind a load balancer, either keep a session on one replica or run one replica. A shared backend alone doesn't fix cancellation, because the cancellation signal is held in the memory of the process running the job.

  • Every caller sees every job. The queue tools have no notion of who submitted a job. In a test with two API keys, the second caller listed the first caller's job with queue_list and cancelled it with queue_cancel. Don't expose the queue tools on a shared multi-tenant server as they are.

  • Job arguments aren't checked up front. queue_submit takes any args dict. A wrong or missing argument only shows up when the job runs, and then counts as a failure that's retried.

  • Job types aren't discoverable. The tool list doesn't say which job types exist. A wrong name gets UNKNOWN_JOB_TYPE with a suggestion listing the available types, but the arguments each one takes are only known if you tell the agent.

  • Results expire. Finished jobs are removed after result_ttl, one hour by default, checked every cleanup_interval of 60 seconds.

  • Testing needs a manual start. TestClient doesn't run startup hooks, so queue workers don't start and submitted jobs stay pending. Call await queue.start() in your test and await queue.stop() after it. The testing example on the docs page uses async with TestClient(server) and reads resp["job_id"]; neither works in 1.2.1, because TestClient isn't a context manager and call_tool returns a list of text content.


[08]

Frequently asked questions

How do I fix an MCP tool timeout?

Stop making the client wait for the whole job. Turn the slow tool into a queue job, so the call that starts it returns a job ID at once and nothing long-running is left to time out. Through Promptise's MCP client, plain tool calls that took 45 seconds and 330 seconds both came back fine, so when you do hit a timeout, check the client you use, any proxy in between, and TimeoutMiddleware on the server. Whatever the limit, a job ID sidesteps it.

How do MCP progress notifications work?

The client sends a progress token with a tool call, and the server sends notifications/progress messages tied to that token while the call runs, each with a number, an optional total and a message. In Promptise, ProgressReporter.report() sends them, and does nothing when the client didn't ask. They show in client UIs; the model doesn't see them. The MCP specification has the details.

What's the difference between BackgroundTasks and an MCP job queue?

BackgroundTasks runs small follow-up work at the end of one tool call, with no ID, status or retries, and in 1.2.1 the caller waits for it. MCPQueue runs jobs on separate workers, gives each a job ID, and adds progress, retries, timeouts and cancellation. Use the queue for anything that takes noticeable time.

Can the MCP job queue use Redis?

Not out of the box in 1.2.1. MCPQueue takes any object that implements the QueueBackend protocol through its backend argument, so you can write a Redis or Postgres backend yourself. The Queue & Background Jobs page shows the methods to implement.

Should the agent wait for a background job?

Usually not. The value of a job is that the conversation can move on, as in Step 3. When the user does want to wait, give the agent a tool that waits on the server side, like wait_for_job in Step 4, instead of letting it poll.


[09]

Where to go next

  • Queue & Background Jobs: every MCPQueue option, priorities, backends and health checks.

  • Resilience Patterns: BackgroundTasks, ProgressReporter, CancellationToken and the circuit breaker.

  • Production Features: caching, rate limits, timeouts and the order to add middleware in.

  • Deployment: running an MCP server over HTTP in production.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and client basics this guide builds on.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server from an existing API, then add jobs for its slow endpoints.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths