Some tools are slow. A quarterly report, a full export or a site crawl can take minutes, and an MCP tool call that takes minutes holds up the agent, the user waiting on it, and sooner or later a client timeout. The fix is the same one web backends use: start the work in the background, hand back a job ID straight away, and let the caller check on it. In this guide you'll build an MCP server in Python whose slow report runs as a background job with progress, retries and cancellation, then put a real agent in front of it. You'll also see what progress notifications, cancellation and fire-and-forget tasks do on a normal tool call, and where the limits are. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed.
How do you handle long-running tasks in an MCP server?
Don't do the slow work inside the tool call. Register it as a job on a queue, so the agent submits it, gets a job ID back in milliseconds, and checks its status or fetches its result whenever it likes. With Promptise Foundry, that's MCPQueue and the @queue.job() decorator:
import asyncio
from promptise.mcp.server import MCPQueue, MCPServer
server = MCPServer("reports")
queue = MCPQueue(server, max_workers=2)
@queue.job(name="generate_report", timeout=120)
async def generate_report(region: str, quarter: str) -> dict:
"""Build the quarterly sales report for one region."""
await asyncio.sleep(8) # the slow part: queries, joins, rendering
return {"region": region, "quarter": quarter, "url": f"https://reports.example.com/{region}-{quarter}.pdf"}
if __name__ == "__main__":
server.run(transport="http", host="127.0.0.1", port=8334)Passing the server to MCPQueue adds five tools to it and starts the workers when the server starts. Submitting the eight-second report returns at once:
['queue_submit', 'queue_status', 'queue_result', 'queue_cancel', 'queue_list']
0.00s {"job_id": "c1d57b4056bae0f0", "status": "pending", "job_type": "generate_report"}The rest of this guide builds this out with progress, retries, cancellation and an agent that uses it.
[02]
How background jobs work in an MCP server
The tool call only puts the job on a queue. A pool of workers inside the server process picks jobs up and runs them, and every later call reads the job's record:
Rendering diagram…
These are the five tools the queue adds. Pass tool_prefix="jobs" to MCPQueue and they become jobs_submit, jobs_status and so on.
Tool | What it does |
|---|---|
queue_submit | Queues a job by job_type with an args dict and an optional priority, and returns its job_id |
queue_status | Returns the status, progress from 0.0 to 1.0, the latest progress message and the attempt count |
queue_result | Returns the result once the job is finished, or the current status while it's still running |
queue_cancel | Cancels a pending or running job |
queue_list | Lists jobs, optionally filtered by status |
A job is pending until a worker takes it, then running, and ends as completed, failed, timeout or cancelled.
[03]
What you need
Python 3.10 or newer.
Promptise Foundry 1.2.1 from PyPI. The server and the queue need nothing else.
An API key for a model provider for the agent steps. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string.
pip install promptise
export OPENAI_API_KEY="sk-..."[04]
Build a report server with background jobs, step by step
The server makes quarterly sales reports, which take a while, and exports invoices to an accounting system that sometimes fails. The complete file, report_server.py, is in Step 4.
Turn the slow function into a job
A job is an ordinary async function. Two extra parameters are filled in by the queue, based on their type annotation: a ProgressReporter that writes progress into the job's record, and a CancellationToken that tells the job when someone has cancelled it.
@queue.job(name="generate_report", timeout=120)
async def generate_report(
region: str,
quarter: str,
progress: ProgressReporter,
cancel: CancellationToken,
) -> dict:
"""Build the quarterly sales report for one region."""
data = SALES[region.lower()]
for done, step in enumerate(STEPS):
cancel.check() # stop here if someone called queue_cancel
await progress.report(done, total=len(STEPS), message=step)
await asyncio.sleep(2) # the slow part: queries, joins, rendering
return {
"region": region,
"quarter": quarter,
"net_revenue": data["revenue"] - data["refunds"],
"orders": data["orders"],
"url": f"https://reports.example.com/{region.lower()}-{quarter.lower()}.pdf",
}progress.report(done, total=4, message=...) stores done / total as a fraction, so status shows 0.25, not 25.
cancel.check() raises when the job has been cancelled. Call it between steps, at points where stopping is safe.
timeout=120 is the ceiling for one attempt. The default is 300 seconds, set by default_timeout on MCPQueue.
max_workers=2 on the queue means at most two jobs run at the same time. The rest wait as pending.
Watch a job from a client
Before an agent gets involved, check the queue with a plain MCP client. Start the server with python report_server.py, then run this:
import asyncio
import json
import time
from promptise.mcp.client import MCPClient
def show(result) -> dict:
return json.loads(result.content[0].text)
async def main():
async with MCPClient(url="http://127.0.0.1:8330/mcp") as client:
print("Tools:", [t.name for t in await client.list_tools()])
start = time.perf_counter()
job = show(await client.call_tool(
"queue_submit",
{"job_type": "generate_report", "args": {"region": "EMEA", "quarter": "Q3"}},
))
print(f"{time.perf_counter() - start:.2f}s submit ->", job)
while True:
await asyncio.sleep(1.5)
status = show(await client.call_tool("queue_status", {"job_id": job["job_id"]}))
print(f"{time.perf_counter() - start:.2f}s status ->", status)
if status["status"] not in ("pending", "running"):
break
result = show(await client.call_tool("queue_result", {"job_id": job["job_id"]}))
print(f"{time.perf_counter() - start:.2f}s result ->", json.dumps(result, indent=2))
asyncio.run(main())Tools: ['queue_submit', 'queue_status', 'queue_result', 'queue_cancel', 'queue_list', 'wait_for_job']
0.00s submit -> {'job_id': '7ff8b16ec4b0f528', 'status': 'pending', 'job_type': 'generate_report'}
1.51s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 1, 'progress_message': 'Loading orders'}
3.01s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}
4.52s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.5, 'attempts': 1, 'progress_message': 'Computing totals'}
…
9.04s status -> {'job_id': '7ff8b16ec4b0f528', 'job_type': 'generate_report', 'status': 'completed', 'priority': 'normal', 'progress': 1.0, 'attempts': 1, 'progress_message': 'Rendering the PDF'}
9.04s result -> {
"job_id": "7ff8b16ec4b0f528",
…
"result": {
"region": "EMEA",
"quarter": "Q3",
"net_revenue": 402450.0,
"orders": 1840,
"url": "https://reports.example.com/emea-q3.pdf"
}
}The submit took no measurable time, the status climbed through the four steps, and the result arrived nine seconds later. wait_for_job in the tool list is a small tool you'll add in Step 4.
Let an agent submit the job and come back later
The model only sees the five queue tools. Nothing in them says which job types exist or what arguments each one takes, so tell the agent in its instructions:
import asyncio
from promptise import build_agent
from promptise.config import HTTPServerSpec
INSTRUCTIONS = """You are a reporting assistant. Reports are slow, so they run as background jobs.
Job types on the reports server:
- generate_report: args {"region": "EMEA" or "Americas", "quarter": "Q1" to "Q4"}
To produce a report, submit it with queue_submit and check it with queue_status.
When the status is completed, fetch it with queue_result. Don't wait for a
slow job: tell the user the job ID and the progress message, and check again
when they ask."""
async def main():
agent = await build_agent(
model="openai:gpt-5-mini",
servers={"reports": HTTPServerSpec(url="http://127.0.0.1:8330/mcp")},
instructions=INSTRUCTIONS,
trace_tools=True,
)
try:
messages = [{"role": "user", "content": "I need the Q3 sales report for EMEA."}]
result = await agent.ainvoke({"messages": messages})
print(">>>", result["messages"][-1].content)
await asyncio.sleep(10) # the user does something else for a while
messages = result["messages"] + [{"role": "user", "content": "Is it ready?"}]
result = await agent.ainvoke({"messages": messages})
print(">>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'EMEA', 'quarter': 'Q3'}, 'priority': 'normal'}
✔ Tool result from queue_submit: {"job_id": "21aa2cc272fed921", "status": "pending", "job_type": "generate_report"}
>>> I've submitted a background job to generate the Q3 sales report for EMEA.
Job ID: 21aa2cc272fed921
Status: pending
…
→ Invoking tool: queue_status with {'job_id': '21aa2cc272fed921'}
✔ Tool result from queue_status: {"job_id": "21aa2cc272fed921", "job_type": "generate_report", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 1, "progress_message": "Rendering the PDF"}
→ Invoking tool: queue_result with {'job_id': '21aa2cc272fed921'}
✔ Tool result from queue_result: {"job_id": "21aa2cc272fed921", …, "result": {"region": "EMEA", "quarter": "Q3", "net_revenue": 402450.0, "orders": 1840, "url": "https://reports.example.com/emea-q3.pdf"}}
>>> Yes — the report is ready.
…
You can download the report here:
https://reports.example.com/emea-q3.pdf
…This is the point of a queue. The first turn ended as soon as the job was queued, so the user got an answer instead of a spinner. The job kept running between turns, and the second turn found it finished. The job ID lives in the conversation history, which is why passing result["messages"] into the next turn matters.
Let the agent wait without polling
Sometimes the user does want to wait. With only the queue tools, an agent waits by calling queue_status again and again, and each check costs a model round trip. With instructions to keep checking until the job finished, and no other tool to wait with, this is how a run went:
→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'EMEA', 'quarter': 'Q3'}, 'priority': 'normal'}
…
→ Invoking tool: queue_status with {'job_id': '64720c23a6330a2e'}
✔ Tool result from queue_status: {"job_id": "64720c23a6330a2e", "job_type": "generate_report", "status": "running", "priority": "normal", "progress": 0.5, "attempts": 1, "progress_message": "Computing totals"}
>>> Your EMEA Q3 report (job_id: 64720c23a6330a2e) is running.
Current status:
- Status: running
- Progress: 0%
- Message: "Loading orders"
…Five status checks in 13 seconds, then it stopped at 50% and handed the decision back to the user. It even reported the first status it had seen, not the latest.
A better way is one tool that waits on the server side and returns when the job finishes or the time is up. It uses the queue's own status() method, so it reads the same record as queue_status. Here it is in the complete server:
import asyncio
import time
from typing import Annotated
from pydantic import Field
from promptise.mcp.server import CancellationToken, MCPQueue, MCPServer, ProgressReporter
server = MCPServer("reports")
queue = MCPQueue(server, max_workers=2)
# A stand-in for your warehouse and your accounting system.
SALES = {
"emea": {"orders": 1840, "revenue": 412_300.0, "refunds": 9_850.0},
"americas": {"orders": 2310, "revenue": 598_120.0, "refunds": 14_200.0},
}
STEPS = ["Loading orders", "Matching refunds", "Computing totals", "Rendering the PDF"]
export_attempts: dict[str, int] = {}
@queue.job(name="generate_report", timeout=120)
async def generate_report(
region: str,
quarter: str,
progress: ProgressReporter,
cancel: CancellationToken,
) -> dict:
"""Build the quarterly sales report for one region."""
data = SALES[region.lower()]
for done, step in enumerate(STEPS):
cancel.check() # stop here if someone called queue_cancel
await progress.report(done, total=len(STEPS), message=step)
await asyncio.sleep(2) # the slow part: queries, joins, rendering
return {
"region": region,
"quarter": quarter,
"net_revenue": data["revenue"] - data["refunds"],
"orders": data["orders"],
"url": f"https://reports.example.com/{region.lower()}-{quarter.lower()}.pdf",
}
@queue.job(name="export_invoices", timeout=60, max_retries=2, backoff_base=2.0)
async def export_invoices(month: str) -> dict:
"""Send one month of invoices to the accounting system."""
export_attempts[month] = export_attempts.get(month, 0) + 1
await asyncio.sleep(1)
if export_attempts[month] == 1:
raise ConnectionError("accounting API answered 503 Service Unavailable")
return {"month": month, "invoices": 214, "file": f"invoices-{month}.csv"}
@server.tool()
async def wait_for_job(
job_id: str,
seconds: Annotated[int, Field(ge=1, le=25, description="How long to wait at most.")] = 20,
) -> dict:
"""Wait for a background job to finish, then return its status.
Use this instead of calling queue_status again and again. It returns as soon as
the job is completed, failed, cancelled or timed out, or when the time is up.
"""
deadline = time.monotonic() + seconds
while True:
status = await queue.status(job_id)
if status["status"] not in ("pending", "running") or time.monotonic() >= deadline:
return status
await asyncio.sleep(0.5)
if __name__ == "__main__":
server.run(transport="http", host="127.0.0.1", port=8330)The 25-second cap keeps each call well inside client timeouts; if the job needs longer, the agent simply calls it again. The agent's instructions now list both job types and say: "When the user wants to wait for the result, call wait_for_job instead of polling queue_status". Asked for two things at once, it ran them side by side:
→ Invoking tool: queue_submit with {'job_type': 'export_invoices', 'args': {'month': '2026-09'}, 'priority': 'normal'}
→ Invoking tool: queue_submit with {'job_type': 'generate_report', 'args': {'region': 'Americas', 'quarter': 'Q3'}, 'priority': 'normal'}
✔ Tool result from queue_submit: {"job_id": "e1ada74daf6a3c9f", "status": "pending", "job_type": "export_invoices"}
✔ Tool result from queue_submit: {"job_id": "d74bcec543b0c071", "status": "pending", "job_type": "generate_report"}
→ Invoking tool: wait_for_job with {'job_id': 'e1ada74daf6a3c9f', 'seconds': 60}
→ Invoking tool: wait_for_job with {'job_id': 'd74bcec543b0c071', 'seconds': 60}
✔ Tool result from wait_for_job: Input validation error: 60 is greater than the maximum of 25
✔ Tool result from wait_for_job: Input validation error: 60 is greater than the maximum of 25
→ Invoking tool: wait_for_job with {'job_id': 'e1ada74daf6a3c9f', 'seconds': 25}
→ Invoking tool: wait_for_job with {'job_id': 'd74bcec543b0c071', 'seconds': 25}
✔ Tool result from wait_for_job: {"job_id": "e1ada74daf6a3c9f", "job_type": "export_invoices", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 2, "error": "Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s."}
✔ Tool result from wait_for_job: {"job_id": "d74bcec543b0c071", "job_type": "generate_report", "status": "completed", "priority": "normal", "progress": 1.0, "attempts": 1, "progress_message": "Rendering the PDF"}
→ Invoking tool: queue_result with {'job_id': 'e1ada74daf6a3c9f'}
→ Invoking tool: queue_result with {'job_id': 'd74bcec543b0c071'}
…
>>> Both background jobs have finished. Summary:
…
- Attempts: 2 (note: attempt 1 failed — "accounting API answered 503 Service Unavailable" — retried and completed)
…
(21.1s)The model asked to wait 60 seconds, the le=25 limit refused it with a clear message, and it tried again with 25. After that, one wait per job was enough, and the export's retry happened without the agent having to do anything. The tool's description does the persuading: with the polling instructions from above and wait_for_job on the server, the agent checked once and then chose wait_for_job by itself.
The model still makes its own choices. In another run it waited for both jobs but fetched only the export's result, and never mentioned the report. If a result must not be dropped, check the job list yourself after the turn rather than relying on the model to.
Retry jobs that fail
export_invoices fails on its first attempt with a ConnectionError, the kind of transient error worth retrying. max_retries=2 allows two more attempts, and backoff_base=2.0 sets the wait between them: backoff_base * 2 ** (attempt - 1), so 2 seconds, then 4. A small client that prints each change of status shows the whole cycle:
0.0s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'pending', 'priority': 'normal', 'progress': 0.0, 'attempts': 0}
0.3s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 1}
1.1s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'pending', 'priority': 'normal', 'progress': 0.0, 'attempts': 1, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
3.3s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'running', 'priority': 'normal', 'progress': 0.0, 'attempts': 2, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
4.1s {'job_id': 'c9a5387b0eda2701', 'job_type': 'export_invoices', 'status': 'completed', 'priority': 'normal', 'progress': 1.0, 'attempts': 2, 'error': 'Attempt 1 failed: accounting API answered 503 Service Unavailable. Retrying in 2.0s.'}
result: {'month': '2026-09', 'invoices': 214, 'file': 'invoices-2026-09.csv'}What to notice:
Between attempts the job is `pending` again, with the failure in error. That error stays on the record after the retry succeeds, so read status to decide whether a job worked, not the presence of error.
Any exception triggers a retry, not only transient ones. A job submitted with a missing argument failed three times with
needs_args() missing 1 required positional argument: 'quarter'before it gave up. Make jobs that can't succeed fail fast, or keep max_retries low.Timeouts aren't retried. A job that runs past its timeout ends as timeout with
Job timed out after 1sin error, on the first attempt.During the backoff, the worker waits too. In 1.2.1 the worker that ran the failed attempt sleeps through the backoff before requeueing, so long backoffs tie up workers. Size max_workers with that in mind.
Cancel a job
queue_cancel marks the job as cancelled and signals its CancellationToken. The job stops the next time it calls cancel.check():
import asyncio
import json
from promptise.mcp.client import MCPClient
def show(result) -> dict:
return json.loads(result.content[0].text)
async def main():
async with MCPClient(url="http://127.0.0.1:8330/mcp") as client:
job = show(await client.call_tool(
"queue_submit", {"job_type": "generate_report", "args": {"region": "Americas", "quarter": "Q3"}}
))
await asyncio.sleep(3)
print("before:", show(await client.call_tool("queue_status", {"job_id": job["job_id"]})))
print("cancel:", show(await client.call_tool("queue_cancel", {"job_id": job["job_id"]})))
await asyncio.sleep(10) # long after the job would have finished
print("later: ", show(await client.call_tool("queue_result", {"job_id": job["job_id"]})))
asyncio.run(main())before: {'job_id': '8a81d9b4c31a23d6', 'job_type': 'generate_report', 'status': 'running', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}
cancel: {'job_id': '8a81d9b4c31a23d6', 'status': 'cancelled'}
later: {'job_id': '8a81d9b4c31a23d6', 'job_type': 'generate_report', 'status': 'cancelled', 'priority': 'normal', 'progress': 0.25, 'attempts': 1, 'progress_message': 'Matching refunds'}The job was halfway through its second step. It stopped at the next cancel.check(), so it never reached "Computing totals", and ten seconds later it's still cancelled with no result.
[05]
Progress and cancellation on a normal tool call
Not everything needs a queue. A tool that takes ten or twenty seconds can stay a normal tool and still tell the client how it's doing. ProgressReporter and CancellationToken work there too, injected with Depends:
@server.tool()
async def crawl_site(pages: int, progress: ProgressReporter = Depends(ProgressReporter)) -> dict:
"""Crawl a number of pages and report progress as it goes."""
for page in range(1, pages + 1):
await asyncio.sleep(0.5)
await progress.report(page, total=pages, message=f"Crawled page {page} of {pages}")
return {"pages": pages}Here, report() sends MCP progress notifications to the client while the call is still running. A client only receives them if it asked for them by sending a progress token with the call. The official MCP Python SDK does that when you pass a progress_callback:
async def on_progress(progress, total, message):
print(f" progress {progress}/{total}: {message}")
async with streamablehttp_client(URL) as (read, write, _):
async with ClientSession(read, write) as session:
await session.initialize()
print("official client, with progress_callback:")
r = await session.call_tool("crawl_site", {"pages": 3}, progress_callback=on_progress)
print(" result:", r.content[0].text)official client, with progress_callback:
progress 1.0/3.0: Crawled page 1 of 3
progress 2.0/3.0: Crawled page 2 of 3
progress 3.0/3.0: Crawled page 3 of 3
result: {"pages": 3}
…
Promptise MCPClient:
result: {"pages": 3} (no progress lines above)Promptise's own MCP client, which a Promptise agent uses, doesn't ask for progress in 1.2.1, so the same calls send nothing to it. The model never sees progress notifications either way; they're for a person watching a client UI. If the agent needs to know how far along something is, use a queue job, whose progress the agent can read with queue_status.
Cancelling a normal tool call works differently from cancelling a job. When the client sends an MCP cancellation, the MCP SDK cancels the running handler outright. The CancellationToken isn't set. A tool that waited on its token logged this when the official client cancelled it:
19:30:54 watch_feed: asyncio task cancelled, token.is_cancelled=FalseThe client got McpError Request cancelled. So on a normal tool, the token won't tell you about a cancellation in 1.2.1. Put any cleanup in a try/finally or catch asyncio.CancelledError, and keep cancel.check() for queue jobs, where it does work.
[06]
Fire-and-forget work with BackgroundTasks
BackgroundTasks is for small jobs after a tool has done its real work, such as sending a receipt or writing an analytics event. You add callables during the tool call and they run once the handler returns:
async def send_receipt(order_id: str) -> None:
await asyncio.sleep(3) # a slow email API
print(f"{time.strftime('%X')} receipt sent for {order_id}", flush=True)
@server.tool()
async def place_order(order_id: str, bg: BackgroundTasks = Depends(BackgroundTasks)) -> dict:
"""Place an order and email the receipt afterwards."""
bg.add(send_receipt, order_id)
print(f"{time.strftime('%X')} handler returned for {order_id}", flush=True)
return {"order_id": order_id, "status": "placed"}Errors in background tasks are logged and never reach the caller. But look at the timing over HTTP. The server printed:
19:30:48 handler returned for A-1001
19:30:51 receipt sent for A-1001And the client calling place_order printed:
19:30:51 place_order answered after 3.01s: {"order_id": "A-1001", "status": "placed"}The handler returned at once, yet the caller waited three seconds for its answer. In 1.2.1, background tasks run before the response goes out, not after it, so they add their full duration to the call. The docs describe the opposite. Until that changes, use BackgroundTasks only for work that takes milliseconds, and make anything slower a queue job.
Here's how to choose:
The work takes | Use | The agent sees |
|---|---|---|
A few seconds | A normal tool | The result |
Up to a minute or so | A normal tool with ProgressReporter | The result; a client UI can show progress |
Minutes, or it may fail and need retries | An MCPQueue job | A job ID now, status and result later |
Milliseconds of follow-up work | BackgroundTasks | The result, after the tasks finish |
[07]
Honest limits
These are true of Promptise Foundry 1.2.1. Each one was checked in a run or in the source.
Jobs live in memory. The only backend that ships is InMemoryQueueBackend. When the server restarts, every job and result is gone: asked about a job submitted before the restart, the server answered JOB_NOT_FOUND. There's no persistence and no Redis backend. You can write your own against the QueueBackend protocol, seven async methods from enqueue to count, but nothing re-runs jobs that were running when a process died.
One queue per process. With two replicas, a job submitted to one is JOB_NOT_FOUND on the other. Behind a load balancer, either keep a session on one replica or run one replica. A shared backend alone doesn't fix cancellation, because the cancellation signal is held in the memory of the process running the job.
Every caller sees every job. The queue tools have no notion of who submitted a job. In a test with two API keys, the second caller listed the first caller's job with queue_list and cancelled it with queue_cancel. Don't expose the queue tools on a shared multi-tenant server as they are.
Job arguments aren't checked up front. queue_submit takes any args dict. A wrong or missing argument only shows up when the job runs, and then counts as a failure that's retried.
Job types aren't discoverable. The tool list doesn't say which job types exist. A wrong name gets UNKNOWN_JOB_TYPE with a suggestion listing the available types, but the arguments each one takes are only known if you tell the agent.
Results expire. Finished jobs are removed after result_ttl, one hour by default, checked every cleanup_interval of 60 seconds.
Testing needs a manual start. TestClient doesn't run startup hooks, so queue workers don't start and submitted jobs stay pending. Call await queue.start() in your test and await queue.stop() after it. The testing example on the docs page uses async with TestClient(server) and reads resp["job_id"]; neither works in 1.2.1, because TestClient isn't a context manager and call_tool returns a list of text content.
[08]
Frequently asked questions
How do I fix an MCP tool timeout?
Stop making the client wait for the whole job. Turn the slow tool into a queue job, so the call that starts it returns a job ID at once and nothing long-running is left to time out. Through Promptise's MCP client, plain tool calls that took 45 seconds and 330 seconds both came back fine, so when you do hit a timeout, check the client you use, any proxy in between, and TimeoutMiddleware on the server. Whatever the limit, a job ID sidesteps it.
How do MCP progress notifications work?
The client sends a progress token with a tool call, and the server sends notifications/progress messages tied to that token while the call runs, each with a number, an optional total and a message. In Promptise, ProgressReporter.report() sends them, and does nothing when the client didn't ask. They show in client UIs; the model doesn't see them. The MCP specification has the details.
What's the difference between BackgroundTasks and an MCP job queue?
BackgroundTasks runs small follow-up work at the end of one tool call, with no ID, status or retries, and in 1.2.1 the caller waits for it. MCPQueue runs jobs on separate workers, gives each a job ID, and adds progress, retries, timeouts and cancellation. Use the queue for anything that takes noticeable time.
Can the MCP job queue use Redis?
Not out of the box in 1.2.1. MCPQueue takes any object that implements the QueueBackend protocol through its backend argument, so you can write a Redis or Postgres backend yourself. The Queue & Background Jobs page shows the methods to implement.
Should the agent wait for a background job?
Usually not. The value of a job is that the conversation can move on, as in Step 3. When the user does want to wait, give the agent a tool that waits on the server side, like wait_for_job in Step 4, instead of letting it poll.
[09]
Where to go next
Queue & Background Jobs: every MCPQueue option, priorities, backends and health checks.
Resilience Patterns: BackgroundTasks, ProgressReporter, CancellationToken and the circuit breaker.
Production Features: caching, rate limits, timeouts and the order to add middleware in.
Deployment: running an MCP server over HTTP in production.
How to Connect MCP Servers to Your AI Agent in Python: the agent and client basics this guide builds on.
OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server from an existing API, then add jobs for its slow endpoints.