Every question your agent answers costs tokens, and every answer depends on one provider staying up. A semantic cache cuts the first cost: when a user asks something they've asked before, in the same or different words, the agent replies from the cache instead of calling the model. Model fallback handles the second: when the primary model fails, a backup answers instead. This guide adds both to a Python support agent with Promptise Foundry, measures the hits, misses, latency and tokens, and simulates a real outage to watch the fallback and circuit breaker work. Every snippet ran against Promptise Foundry 1.2.1, and the guide is plain about the places where 1.2.1 falls short.
How do you add an LLM semantic cache and fallback to an agent?
Pass a SemanticCache as cache= and a FallbackChain as model= when you build the agent, and tell each request who is asking:
import asyncio
import time
from promptise import CallerContext, FallbackChain, SemanticCache, build_agent
from acme_support import INSTRUCTIONS
async def main():
cache = SemanticCache()
cache.warmup() # load the local embedding model now, not on the first question
agent = await build_agent(
model=FallbackChain(["openai:gpt-5-mini", "openai:gpt-4.1-mini"]),
servers={},
instructions=INSTRUCTIONS,
cache=cache,
)
alice = CallerContext(user_id="alice")
try:
for question in ["How do I reset my password?", "How can I reset my password?"]:
start = time.perf_counter()
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
print(f"{time.perf_counter() - start:5.2f}s {result['messages'][-1].content}")
finally:
await agent.shutdown()
asyncio.run(main())[promptise] No tools discovered from MCP servers; agent will run without tools.
5.47s Go to https://acme.example/reset and follow the instructions to reset your password. The reset link you receive expires after 30 minutes.
0.06s Go to https://acme.example/reset and follow the instructions to reset your password. The reset link you receive expires after 30 minutes.The first line is Promptise noting that this agent has no tools, which is fine for a support bot that answers from its instructions. The second question uses different words, but it means the same thing, so the cache answered it in 60 milliseconds without calling the model. If gpt-5-mini had failed, gpt-4.1-mini would have answered instead. Two details matter: with the default settings, the cache only works for requests that carry a CallerContext with a user_id, and in 1.2.1 FallbackChain only works for agents without tools. Both are covered below.
[02]
How semantic caching and fallback fit together
A normal cache needs the exact same key. A semantic cache turns each question into an embedding, a list of numbers that captures its meaning, and compares it with the questions already cached. If the closest one is similar enough, by cosine similarity, its answer comes back. Only on a miss does the request reach the model, and that's where the fallback chain sits:
Rendering diagram…
A similar question isn't enough on its own. A cached answer is only reused when all of these match:
Part of the cache key | What it prevents |
|---|---|
Scope, such as user:alice | One user seeing another user's answers |
Question embedding, above similarity_threshold | Answering a different question |
Model name | Serving one model's answer for another |
Hash of the instructions | Stale answers after you change the prompt |
Context fingerprint: memory results and message count | Stale answers after memory changes |
The cache check runs after input guardrails and memory search, and output guardrails run again on every cached answer before it's returned. Semantic Cache in the docs has the full list.
[03]
What you need
Python 3.10 or newer. This guide used Python 3.12.
Promptise Foundry plus the two packages the cache needs. numpy does the similarity maths, and sentence-transformers runs the default embedding model on your machine. It pulls in PyTorch, which is a large download.
An OpenAI API key. The default embedding model, all-MiniLM-L6-v2, downloads from Hugging Face on first use, about 90 MB.
pip install "promptise==1.2.1" numpy sentence-transformers
export OPENAI_API_KEY="sk-..."[04]
Cut LLM costs with a semantic cache, step by step
The agent is a support assistant for a made-up product, Acme Cloud. Its facts live in the instructions, so every answer comes straight from the model, which is exactly the kind of traffic a cache helps with:
INSTRUCTIONS = """You are the support assistant for Acme Cloud. Answer in at most two sentences.
Facts you can rely on:
- Reset a password at https://acme.example/reset. The reset link expires after 30 minutes.
- The Free plan includes 3 projects and 1 GB of storage.
- The Pro plan includes 50 projects and 100 GB of storage, for $20 a month.
- Human support is available Monday to Friday, 8:00 to 18:00 CET."""Check how similar your questions really are
The threshold decides everything. Too low and the cache answers questions it shouldn't; too high and it never hits. Before you pick one, score a few real question pairs with the same embedding providers the cache uses:
import asyncio
import os
from promptise import LocalEmbeddingProvider, OpenAIEmbeddingProvider
# (question already in the cache, new question)
PAIRS = [
("How do I reset my password?", "How do I reset my password?"),
("How do I reset my password?", "how do i reset my password"),
("How do I reset my password?", "How can I reset my password?"),
("How do I reset my password?", "What's the way to reset my password?"),
("How do I reset my password?", "I forgot my password, what now?"),
("How do I reset my password?", "How do I reset my API key?"),
("How many projects does the Free plan include?", "How many projects does the Pro plan include?"),
]
async def similarities(provider) -> list[float]:
vectors = await provider.embed([text for pair in PAIRS for text in pair])
# Both providers return unit-length vectors, so the dot product is the cosine similarity.
return [sum(a * b for a, b in zip(vectors[i], vectors[i + 1])) for i in range(0, len(vectors), 2)]
async def main():
local = await similarities(LocalEmbeddingProvider())
openai = await similarities(OpenAIEmbeddingProvider(api_key=os.environ["OPENAI_API_KEY"]))
print(f"{'Cached question':<46} {'New question':<46} {'local':>6} {'openai':>7}")
for (cached, new), a, b in zip(PAIRS, local, openai):
print(f"{cached:<46} {new:<46} {a:6.3f} {b:7.3f}")
asyncio.run(main())Cached question New question local openai
How do I reset my password? How do I reset my password? 1.000 1.000
How do I reset my password? how do i reset my password 0.994 0.908
How do I reset my password? How can I reset my password? 0.992 0.961
How do I reset my password? What's the way to reset my password? 0.966 0.910
How do I reset my password? I forgot my password, what now? 0.795 0.714
How do I reset my password? How do I reset my API key? 0.525 0.592
How many projects does the Free plan include? How many projects does the Pro plan include? 0.793 0.832With the default threshold of 0.92, the local model treats the first four as the same question and the last three as different, which is right. "I forgot my password, what now?" means the same thing but scores 0.795, so it will miss. That's the safe kind of mistake: it costs a model call, not a wrong answer.
The two columns also show that scores aren't portable between embedding models. OpenAI's text-embedding-3-small scores the lowercase version at 0.908, below the default threshold. Pick the threshold for the embedding model you actually use. Here's the same lowercase repeat through a cached agent that uses OpenAI embeddings, first at the default threshold and then at 0.88:
cache = SemanticCache(
embedding=OpenAIEmbeddingProvider(model="text-embedding-3-small", api_key=os.environ["OPENAI_API_KEY"]),
similarity_threshold=threshold,
)…
threshold 0.92: MISS How do I reset my password?
threshold 0.92: MISS how do i reset my password
…
threshold 0.88: MISS How do I reset my password?
threshold 0.88: HIT how do i reset my passwordOpenAI embeddings skip the PyTorch install and the local model, but every cache lookup becomes an API call of its own, which adds latency to hits and misses alike.
Add the cache and measure it
Now run eight questions through a cached agent. LangChain's get_usage_metadata_callback counts the tokens of every model call made inside the with block, so a real cache hit shows zero:
import asyncio
import time
from langchain_core.callbacks import get_usage_metadata_callback
from promptise import CallerContext, SemanticCache, build_agent
from acme_support import INSTRUCTIONS
QUESTIONS = [
"How do I reset my password?",
"How do I reset my password?",
"how do i reset my password",
"What's the way to reset my password?",
"I forgot my password, what now?",
"How many projects does the Free plan include?",
"How many projects does the Pro plan include?",
"How many projects are included in the Free plan?",
]
async def main():
cache = SemanticCache() # local embeddings, threshold 0.92, one hour TTL, per-user scope
cache.warmup()
agent = await build_agent(
model="openai:gpt-5-mini",
servers={},
instructions=INSTRUCTIONS,
cache=cache,
)
alice = CallerContext(user_id="alice")
total_tokens = 0
try:
for question in QUESTIONS:
hits_before = (await cache.stats()).hits
start = time.perf_counter()
# Counts the tokens of every model call made inside this block.
with get_usage_metadata_callback() as usage:
await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
elapsed_ms = (time.perf_counter() - start) * 1000
hit = (await cache.stats()).hits > hits_before
tokens = sum(u["total_tokens"] for u in usage.usage_metadata.values())
total_tokens += tokens
print(f"{'HIT ' if hit else 'MISS'} {elapsed_ms:7.0f} ms {tokens:6} tokens {question}")
stats = await cache.stats()
print(f"\n{stats} hit rate {stats.hit_rate:.0%}")
print(f"Tokens used: {total_tokens}")
finally:
await agent.shutdown()
asyncio.run(main())…
MISS 2990 ms 286 tokens How do I reset my password?
HIT 17 ms 0 tokens How do I reset my password?
HIT 620 ms 0 tokens how do i reset my password
HIT 57 ms 0 tokens What's the way to reset my password?
MISS 3185 ms 369 tokens I forgot my password, what now?
MISS 1262 ms 136 tokens How many projects does the Free plan include?
MISS 1854 ms 200 tokens How many projects does the Pro plan include?
HIT 19 ms 0 tokens How many projects are included in the Free plan?
CacheStats(hits=4, misses=4, stores=4, evictions=0) hit rate 50%
Tokens used: 991Half the questions cost nothing and came back in well under a second, most in under 60 ms, against one to three seconds for a model call. The Pro plan question missed even though it looks like the Free plan one, so nobody was told Pro has 3 projects. Your own hit rate depends entirely on how repetitive your traffic is. Measure it with cache.stats() on real questions before you count on savings.
Keep each user's answers to themselves
A cached answer can contain anything the agent said to that user. That's why the default scope is per_user: each user_id gets its own partition, and a request without a caller isn't cached at all. Here's the part of isolation.py that checks it:
async def ask(agent, cache, who: str, caller: CallerContext | None) -> None:
hits_before = (await cache.stats()).hits
await agent.ainvoke(QUESTION, caller=caller)
print(f"{'HIT ' if (await cache.stats()).hits > hits_before else 'MISS'} {who}") alice, bob = CallerContext(user_id="alice"), CallerContext(user_id="bob")
await ask(agent, cache, "alice", alice)
await ask(agent, cache, "alice again", alice)
await ask(agent, cache, "bob", bob)
await ask(agent, cache, "no caller", None)
await ask(agent, cache, "no caller again", None)
print("purged entries for alice:", await cache.purge_user("alice"))
await ask(agent, cache, "alice after purge", alice)scope='per_user' (the default)
…
MISS alice
HIT alice again
MISS bob
MISS no caller
MISS no caller again
purged entries for alice: 1
MISS alice after purgeBob asked exactly what Alice asked and still got a fresh answer. purge_user removes everything cached for one user, which is what you call when someone asks to have their data erased. If your callers carry a tenant_id, it becomes part of the scope too, so two tenants with the same user ID never share entries; pass the same tenant_id to purge_user.
For finer isolation, scope="per_session" keys the cache on a session_id in the caller's metadata:
cache = SemanticCache(scope="per_session")
agent = await build_agent(model="openai:gpt-5-mini", servers={}, instructions=INSTRUCTIONS, cache=cache)
morning = CallerContext(user_id="alice", metadata={"session_id": "chat-1"})
evening = CallerContext(user_id="alice", metadata={"session_id": "chat-2"})scope='per_session'
…
MISS alice, chat-1
MISS alice, chat-2
HIT alice, chat-1 againThe third scope, shared, gives every caller one cache and works without a CallerContext. Use it only when answers never depend on who's asking, such as a public FAQ, and pass shared_data_acknowledged=True, or Promptise logs a warning.
Expire answers that go stale
Every entry has a time to live. default_ttl sets it in seconds, one hour by default. For questions about things that change, ttl_patterns maps a regular expression to a shorter TTL. Patterns are matched against the lowercased question, so write them in lowercase. This runs without a model, using the cache's own store and check:
import asyncio
from promptise import CallerContext, SemanticCache
async def main():
cache = SemanticCache(
default_ttl=3600, # most answers live for an hour
ttl_patterns={r"status|today|price": 2}, # time-sensitive questions expire after 2 seconds
)
alice = CallerContext(user_id="alice")
questions = ["How do I reset my password?", "Is there an outage today?", "What's the service status?"]
for question in questions:
await cache.store(question, "a stored answer", {"messages": []}, caller=alice)
await asyncio.sleep(3)
for question in questions:
entry = await cache.check(question, caller=alice)
print(f"{'HIT ' if entry else 'MISS'} {question}")
asyncio.run(main())HIT How do I reset my password?
MISS Is there an outage today?
MISS What's the service status?In production you'd use something like 60 seconds rather than 2. The cache also misses on its own when you change the instructions or the model, because both are part of the key.
[05]
Survive LLM provider outages with model fallback, step by step
FallbackChain is a chat model that wraps several others. It tries them in order; if one raises an error or exceeds its timeout, the next one gets the same request. Each model has its own circuit breaker: after failure_threshold consecutive failures, the chain skips that model for recovery_timeout seconds, then lets one test request through.
Simulate an outage you can switch on and off
You can't take OpenAI down on demand, so this small server stands in for it. While it's down it returns HTTP 503 to every request, which is what a provider outage looks like to your code. While it's up it forwards requests to the real OpenAI API:
"""A stand-in for your LLM provider that you can switch off and on.
While it's "down", every request gets HTTP 503, like a provider outage.
While it's "up", requests are forwarded to the real OpenAI API.
"""
import httpx
import uvicorn
from starlette.applications import Starlette
from starlette.requests import Request
from starlette.responses import JSONResponse, Response
from starlette.routing import Route
state = {"up": False, "requests": 0}
async def status(request: Request) -> JSONResponse:
return JSONResponse(state)
async def switch(request: Request) -> JSONResponse:
state["up"] = request.path_params["mode"] == "up"
return JSONResponse(state)
async def provider(request: Request) -> Response:
state["requests"] += 1
if not state["up"]:
return JSONResponse(
{"error": {"message": "Service unavailable (simulated outage)", "type": "server_error"}},
status_code=503,
)
async with httpx.AsyncClient(timeout=120) as client:
upstream = await client.post(
f"https://api.openai.com/v1/{request.path_params['path']}",
content=await request.body(),
headers={
"authorization": request.headers["authorization"],
"content-type": "application/json",
},
)
return Response(upstream.content, upstream.status_code, media_type="application/json")
app = Starlette(
routes=[
Route("/admin/status", status, methods=["GET"]),
Route("/admin/{mode:str}", switch, methods=["POST"]),
Route("/v1/{path:path}", provider, methods=["POST"]),
]
)
if __name__ == "__main__":
uvicorn.run(app, host="127.0.0.1", port=8190, log_level="warning")Start it in its own terminal with python outage_proxy.py. Everything it needs comes with Promptise.
Put a fallback chain in front of the model
The primary is gpt-5-mini, pointed at the outage simulator with Model(..., endpoint=...). The backup talks to OpenAI directly. A low threshold and a short recovery time make the circuit breaker easy to watch:
import asyncio
import time
import httpx
from promptise import FallbackChain, Model, build_agent
from acme_support import INSTRUCTIONS
PROXY = "http://127.0.0.1:8190"
def primary_requests() -> int:
return httpx.get(f"{PROXY}/admin/status").json()["requests"]
async def main():
chain = FallbackChain(
[
# The primary goes through the outage simulator. max_retries=0 stops the
# OpenAI SDK from retrying on its own before the chain can fall back.
Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
"openai:gpt-4.1-mini",
],
timeout_per_model=20,
failure_threshold=2,
recovery_timeout=10,
on_fallback=lambda failed, backup, error: print(f" fallback: {failed} -> {backup} ({error})"),
)
agent = await build_agent(model=chain, servers={}, instructions=INSTRUCTIONS)
async def ask(question: str) -> None:
before, start = primary_requests(), time.perf_counter()
await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
elapsed = time.perf_counter() - start
state = chain.get_chain_status()[0]
print(
f"served by {chain.model_name:<20} {elapsed:5.2f}s "
f"primary tried {primary_requests() - before}x primary circuit: {state['state']}"
)
try:
httpx.post(f"{PROXY}/admin/down")
print("Primary is down")
await ask("How do I reset my password?")
await ask("How many projects does the Free plan include?")
await ask("When is human support available?")
print(chain.get_chain_status())
httpx.post(f"{PROXY}/admin/up")
print("\nPrimary is back. Waiting for the recovery timeout...")
await asyncio.sleep(10)
await ask("How much does the Pro plan cost?")
await ask("How much storage does the Pro plan include?")
finally:
await agent.shutdown()
asyncio.run(main())…
Primary is down
fallback: openai:gpt-5-mini -> openai:gpt-4.1-mini (Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}})
served by openai:gpt-4.1-mini 1.24s primary tried 1x primary circuit: closed
fallback: openai:gpt-5-mini -> openai:gpt-4.1-mini (Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}})
served by openai:gpt-4.1-mini 0.56s primary tried 1x primary circuit: open
served by openai:gpt-4.1-mini 0.58s primary tried 0x primary circuit: open
[{'model_id': 'openai:gpt-5-mini', 'state': 'open', 'failures': 2, 'is_primary': True}, {'model_id': 'openai:gpt-4.1-mini', 'state': 'closed', 'failures': 0, 'is_primary': False}]
Primary is back. Waiting for the recovery timeout...
served by openai:gpt-5-mini 1.97s primary tried 1x primary circuit: closed
served by openai:gpt-5-mini 1.45s primary tried 1x primary circuit: closedRead it from the top. The first two requests tried the primary, got a 503, and were answered by the backup. After the second failure the circuit opened, so the third request didn't touch the primary at all. Once the recovery timeout had passed, the next request was the test: the primary answered, the circuit closed, and traffic went back to it. on_fallback is the place to send an alert, and get_chain_status() is what a health check would report.
What counts as a failure? Any exception from the model, plus a timeout when you set timeout_per_model or global_timeout. That includes rate limits and server errors, but also errors that a second model won't fix, such as a request that's too long for the context window.
Turn off the SDK's own retries
The OpenAI client retries failed requests by itself before it gives up, and the fallback can only start after that. retries_demo.py runs the same outage twice, once with the SDK's default and once with extra={"max_retries": 0} on the primary:
for label, extra in [("SDK default retries", {}), ("max_retries=0", {"max_retries": 0})]:
primary = Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra=extra)
chain = FallbackChain([primary, "openai:gpt-4.1-mini"])…
SDK default retries 2.01s primary tried 3x served by openai:gpt-4.1-mini
…
max_retries=0 0.83s primary tried 1x served by openai:gpt-4.1-miniWith retries on, every request during an outage hits the dead provider three times before falling back. Turn retries off on every model except the last one in the chain, where a retry is your final chance.
Let the cache answer when everything is down
The cache sits in front of the chain, so answers it already holds survive even a total outage. In outage_cache.py, both models go through the simulator. One question is asked while it's up, then everything goes down:
chain = FallbackChain(
[
Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
Model("gpt-4.1-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
]
)
cache = SemanticCache()
cache.warmup()
agent = await build_agent(model=chain, servers={}, instructions=INSTRUCTIONS, cache=cache)…
Agent: Go to https://acme.example/reset, enter your account email, and follow the instructions sent to you. The reset link expires after 30 minutes, so use it promptly.
Every model is down
Agent: Go to https://acme.example/reset, enter your account email, and follow the instructions sent to you. The reset link expires after 30 minutes, so use it promptly.
GraphExecutionError: Graph 'react' failed at node 'reason': RuntimeError: All 2 models in FallbackChain failed.
openai:gpt-5-mini: OpenAIAPIError: Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}}
openai:gpt-4.1-mini: OpenAIAPIError: Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}}The paraphrased password question still got an answer. The new question failed with an error that names every model and why it failed, which is what you want in your logs. Catch it and show the user a clear message.
[06]
Agents with tools: what changes
Most agents call tools, and both features behave differently there. This is where you need to be most careful.
The cache replays tool-using turns
Promptise caches a turn whether or not it called tools. On a hit, no tool runs; the agent returns the old answer. The tool names are saved with the entry, but nothing in 1.2.1 acts on them: invalidate_on_write=True is the default and the docs say a write tool evicts the cache, yet the agent never calls the method that does it. Here's a ticket agent with a read tool and a write tool, run with the default cache:
for question in [
"How many open tickets do I have?",
"Open a ticket: the invoice PDF is blank.",
"How many open tickets do I have?",
"Open a ticket: the invoice PDF is blank.",
]:
print(f"\n>>> {question}")
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
print("Agent:", result["messages"][-1].content)
print("Tickets really open:", len(json.loads(DB.read_text())))>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Tickets really open: 2
>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-103", "subject": "Invoice PDF is blank"}
Agent: Ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3
>>> How many open tickets do I have?
Agent: You currently have 2 open support tickets.
Tickets really open: 3
>>> Open a ticket: the invoice PDF is blank.
Agent: Ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3Two problems in four lines. The count is stale, because the write didn't clear the cache. Worse, the second "open a ticket" request never reached the tool, yet the agent said it did. For a tool that sends money or emails, that's a real bug in your product.
Until Promptise wires this up, clear the caller's cache yourself whenever a turn used a write tool. invalidate_for_write is public and removes every entry in that caller's scope, including the one just stored:
WRITE_TOOLS = {"open_ticket"}
alice = CallerContext(user_id="alice")
async def ask(agent, cache: SemanticCache, question: str, caller: CallerContext) -> str:
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=caller)
tools_called = {call["name"] for message in result["messages"] for call in getattr(message, "tool_calls", None) or []}
if tools_called & WRITE_TOOLS:
# Something changed: drop this caller's cached answers, including the one just stored.
await cache.invalidate_for_write(", ".join(sorted(tools_called & WRITE_TOOLS)), caller=caller)
return result["messages"][-1].content>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Tickets really open: 2
>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-103", "subject": "Invoice PDF is blank"}
Agent: Done — ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3
>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 3
Agent: You have 3 open support tickets.
Tickets really open: 3
>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-104", "subject": "Invoice PDF is blank"}
Agent: I've opened ticket T-104 for "Invoice PDF is blank."
Tickets really open: 4Both answers are right now. This only covers writes made through this agent; if data changes elsewhere, rely on a short TTL. The simplest rule is often the best one: put the cache on agents that answer questions, and leave it off agents that take actions.
FallbackChain doesn't work with tools in 1.2.1
To call tools, an agent binds them to the model. FallbackChain doesn't implement that yet, so any agent with tools fails on its first request:
agent = await build_agent(
model=FallbackChain(["openai:gpt-5-mini", "openai:gpt-4.1-mini"]),
servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
)GraphExecutionError: Graph 'react' failed at node 'reason': NotImplementedError:build_agent also accepts a LangChain runnable as the model, and LangChain's own with_fallbacks does support tool calling. In this test the primary's model name is deliberately wrong, so it fails every time:
model = ChatOpenAI(model="gpt-5-mini-does-not-exist", max_retries=0).with_fallbacks(
[ChatOpenAI(model="gpt-4.1-mini", max_retries=0)]
)
agent = await build_agent(
model=model,
servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
instructions="You are a support assistant. Use your tools. Answer in one sentence.",
trace_tools=True,
)→ Invoking tool: count_open_tickets with {'payload': None}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Answered by: gpt-4.1-mini-2025-04-14The backup answered and the tool call worked. What you give up is the circuit breaker: every request tries the primary first, so keep max_retries=0 and a timeout on it to make failures fast.
[07]
Honest limits
The cache saves real money on repeated questions, but a few cases need care. The first one is a bug in 1.2.1 that can give a user a wrong answer.
Multi-turn conversations can get the wrong cached answer. The cache embeds only the last message, and the context fingerprint counts the messages without looking at what they say. cache_multi_turn=False is the default, but 1.2.1 never reads it. So two conversations of the same length that end in the same follow-up collide, as multi_turn.py shows:
…
Tell me about the city of Paris. / What river runs through it?
Agent: The River Seine runs through Paris.
CacheStats(hits=0, misses=1, stores=1, evictions=0)
Tell me about the city of London. / What river runs through it?
Agent: The River Seine runs through Paris.
CacheStats(hits=1, misses=1, stores=1, evictions=0)The rest, in short:
Chat needs `scope="per_session"`. Then one conversation can't answer from another. Or cache only single-question requests.
Personalised answers are only safe per user. Never use scope="shared" for an agent that sees accounts, orders or anything private.
Time-sensitive answers need a short TTL. Prices, statuses and "today" questions belong in ttl_patterns, or out of the cache.
Tool calls with side effects are replayed, not run. Clear the cache after writes, as shown above, or keep the cache off agents that act.
The in-memory backend lives in one process. Each worker has its own cache, and a restart empties it. backend="redis" shares it across workers and survives restarts; it wasn't tested for this guide.
`CacheStats` can count one request twice. When a similar question is found but the context doesn't match, 1.2.1 counts both a hit and a miss. Treat the hit rate as approximate.
[08]
Frequently asked questions
What is semantic caching for LLMs?
A semantic cache stores model answers keyed by the meaning of the question rather than its exact text. Each question becomes an embedding, and a new question that's close enough to a cached one, by cosine similarity, gets the cached answer without a model call. That catches rephrasings and typos that an exact-match cache would miss.
How much does a semantic cache reduce LLM API costs?
Every hit saves a whole model call, so the saving equals your hit rate. In this guide's test, four of eight questions were hits and cost zero tokens. Real traffic varies a lot: FAQ-style support repeats itself, while open-ended chat rarely does. Run the cache on real questions and read cache.stats() before you estimate savings.
What similarity threshold should I use?
Start with the default 0.92 and test it on pairs from your own logs, as in Step 1. Raise it if you see wrong answers; lower it carefully if obvious rephrasings miss. Thresholds don't carry over between embedding models: the same lowercase question scored 0.994 with the local model and 0.908 with OpenAI's.
How does LLM fallback work with LangChain?
LangChain models have a with_fallbacks method that tries a list of backup runnables when the first one raises an error, and it supports tool calling. Promptise's FallbackChain adds per-model timeouts and a circuit breaker on top, but in 1.2.1 it only works for agents without tools. Both can be passed to build_agent(model=...).
Is a semantic cache safe when many users share one agent?
With the default per_user scope, yes: each user ID has its own partition, a request without a CallerContext isn't cached, and output guardrails run on every cached answer. The risks are stale or replayed answers, not leaks between users, and the limits above cover them.
[09]
Where to go next
Semantic Cache: every option, the Redis backend and custom embedding providers.
Model Fallback: FallbackChain options and circuit breaker states.
Models & Providers: Model and the provider strings for your backups.
Building Multi-User Systems: CallerContext, tenants and isolation across Promptise.
Events & Notifications: get alerted when requests fail.
How to Connect MCP Servers to Your AI Agent in Python: give the agent real tools.