PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

LLM Semantic Cache and Model Fallback in Python

Reduce LLM API costs with a semantic cache and survive provider outages with model fallback. Real hits, tokens and a simulated outage, in Python.

Level
Intermediate
Reading time
23 min
Published
Oct 10, 2026
By
Promptise Team
  • Semantic Cache
  • LLM Fallback
  • LLM Cost
  • AI Agents
  • Python
  • Promptise Foundry

Every question your agent answers costs tokens, and every answer depends on one provider staying up. A semantic cache cuts the first cost: when a user asks something they've asked before, in the same or different words, the agent replies from the cache instead of calling the model. Model fallback handles the second: when the primary model fails, a backup answers instead. This guide adds both to a Python support agent with Promptise Foundry, measures the hits, misses, latency and tokens, and simulates a real outage to watch the fallback and circuit breaker work. Every snippet ran against Promptise Foundry 1.2.1, and the guide is plain about the places where 1.2.1 falls short.

[01]

How do you add an LLM semantic cache and fallback to an agent?

Pass a SemanticCache as cache= and a FallbackChain as model= when you build the agent, and tell each request who is asking:

Pythonquickstart.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
import asyncio
import time

from promptise import CallerContext, FallbackChain, SemanticCache, build_agent

from acme_support import INSTRUCTIONS


async def main():
    cache = SemanticCache()
    cache.warmup()  # load the local embedding model now, not on the first question
    agent = await build_agent(
        model=FallbackChain(["openai:gpt-5-mini", "openai:gpt-4.1-mini"]),
        servers={},
        instructions=INSTRUCTIONS,
        cache=cache,
    )
    alice = CallerContext(user_id="alice")
    try:
        for question in ["How do I reset my password?", "How can I reset my password?"]:
            start = time.perf_counter()
            result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
            print(f"{time.perf_counter() - start:5.2f}s  {result['messages'][-1].content}")
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
[promptise] No tools discovered from MCP servers; agent will run without tools.
 5.47s  Go to https://acme.example/reset and follow the instructions to reset your password. The reset link you receive expires after 30 minutes.
 0.06s  Go to https://acme.example/reset and follow the instructions to reset your password. The reset link you receive expires after 30 minutes.

The first line is Promptise noting that this agent has no tools, which is fine for a support bot that answers from its instructions. The second question uses different words, but it means the same thing, so the cache answered it in 60 milliseconds without calling the model. If gpt-5-mini had failed, gpt-4.1-mini would have answered instead. Two details matter: with the default settings, the cache only works for requests that carry a CallerContext with a user_id, and in 1.2.1 FallbackChain only works for agents without tools. Both are covered below.


[02]

How semantic caching and fallback fit together

A normal cache needs the exact same key. A semantic cache turns each question into an embedding, a list of numbers that captures its meaning, and compares it with the questions already cached. If the closest one is similar enough, by cosine similarity, its answer comes back. Only on a miss does the request reach the model, and that's where the fallback chain sits:

Rendering diagram…

A similar question isn't enough on its own. A cached answer is only reused when all of these match:

Part of the cache key

What it prevents

Scope, such as user:alice

One user seeing another user's answers

Question embedding, above similarity_threshold

Answering a different question

Model name

Serving one model's answer for another

Hash of the instructions

Stale answers after you change the prompt

Context fingerprint: memory results and message count

Stale answers after memory changes

The cache check runs after input guardrails and memory search, and output guardrails run again on every cached answer before it's returned. Semantic Cache in the docs has the full list.


[03]

What you need

  • Python 3.10 or newer. This guide used Python 3.12.

  • Promptise Foundry plus the two packages the cache needs. numpy does the similarity maths, and sentence-transformers runs the default embedding model on your machine. It pulls in PyTorch, which is a large download.

  • An OpenAI API key. The default embedding model, all-MiniLM-L6-v2, downloads from Hugging Face on first use, about 90 MB.

>_Terminal
pip install "promptise==1.2.1" numpy sentence-transformers
export OPENAI_API_KEY="sk-..."
Warning

Without sentence-transformers, the cache doesn't raise an error. It logs Cache: embedding failed, skipping cache check on every request and calls the model each time. Call cache.warmup() at startup: it loads the model right away and raises an ImportError if the package is missing.


[04]

Cut LLM costs with a semantic cache, step by step

The agent is a support assistant for a made-up product, Acme Cloud. Its facts live in the instructions, so every answer comes straight from the model, which is exactly the kind of traffic a cache helps with:

Pythonacme_support.py
1
2
3
4
5
6
7
INSTRUCTIONS = """You are the support assistant for Acme Cloud. Answer in at most two sentences.

Facts you can rely on:
- Reset a password at https://acme.example/reset. The reset link expires after 30 minutes.
- The Free plan includes 3 projects and 1 GB of storage.
- The Pro plan includes 50 projects and 100 GB of storage, for $20 a month.
- Human support is available Monday to Friday, 8:00 to 18:00 CET."""
Step 01

Check how similar your questions really are

The threshold decides everything. Too low and the cache answers questions it shouldn't; too high and it never hits. Before you pick one, score a few real question pairs with the same embedding providers the cache uses:

Pythonsimilarity.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
import asyncio
import os

from promptise import LocalEmbeddingProvider, OpenAIEmbeddingProvider

# (question already in the cache, new question)
PAIRS = [
    ("How do I reset my password?", "How do I reset my password?"),
    ("How do I reset my password?", "how do i reset my password"),
    ("How do I reset my password?", "How can I reset my password?"),
    ("How do I reset my password?", "What's the way to reset my password?"),
    ("How do I reset my password?", "I forgot my password, what now?"),
    ("How do I reset my password?", "How do I reset my API key?"),
    ("How many projects does the Free plan include?", "How many projects does the Pro plan include?"),
]


async def similarities(provider) -> list[float]:
    vectors = await provider.embed([text for pair in PAIRS for text in pair])
    # Both providers return unit-length vectors, so the dot product is the cosine similarity.
    return [sum(a * b for a, b in zip(vectors[i], vectors[i + 1])) for i in range(0, len(vectors), 2)]


async def main():
    local = await similarities(LocalEmbeddingProvider())
    openai = await similarities(OpenAIEmbeddingProvider(api_key=os.environ["OPENAI_API_KEY"]))
    print(f"{'Cached question':<46} {'New question':<46} {'local':>6} {'openai':>7}")
    for (cached, new), a, b in zip(PAIRS, local, openai):
        print(f"{cached:<46} {new:<46} {a:6.3f} {b:7.3f}")


asyncio.run(main())
Output
Cached question                                New question                                    local  openai
How do I reset my password?                    How do I reset my password?                     1.000   1.000
How do I reset my password?                    how do i reset my password                      0.994   0.908
How do I reset my password?                    How can I reset my password?                    0.992   0.961
How do I reset my password?                    What's the way to reset my password?            0.966   0.910
How do I reset my password?                    I forgot my password, what now?                 0.795   0.714
How do I reset my password?                    How do I reset my API key?                      0.525   0.592
How many projects does the Free plan include?  How many projects does the Pro plan include?    0.793   0.832

With the default threshold of 0.92, the local model treats the first four as the same question and the last three as different, which is right. "I forgot my password, what now?" means the same thing but scores 0.795, so it will miss. That's the safe kind of mistake: it costs a model call, not a wrong answer.

The two columns also show that scores aren't portable between embedding models. OpenAI's text-embedding-3-small scores the lowercase version at 0.908, below the default threshold. Pick the threshold for the embedding model you actually use. Here's the same lowercase repeat through a cached agent that uses OpenAI embeddings, first at the default threshold and then at 0.88:

Pythonopenai_embeddings.py
        cache = SemanticCache(
            embedding=OpenAIEmbeddingProvider(model="text-embedding-3-small", api_key=os.environ["OPENAI_API_KEY"]),
            similarity_threshold=threshold,
        )
Output
…
threshold 0.92: MISS  How do I reset my password?
threshold 0.92: MISS  how do i reset my password
…
threshold 0.88: MISS  How do I reset my password?
threshold 0.88: HIT   how do i reset my password

OpenAI embeddings skip the PyTorch install and the local model, but every cache lookup becomes an API call of its own, which adds latency to hits and misses alike.

Step 02

Add the cache and measure it

Now run eight questions through a cached agent. LangChain's get_usage_metadata_callback counts the tokens of every model call made inside the with block, so a real cache hit shows zero:

Pythoncache_demo.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
import asyncio
import time

from langchain_core.callbacks import get_usage_metadata_callback

from promptise import CallerContext, SemanticCache, build_agent

from acme_support import INSTRUCTIONS

QUESTIONS = [
    "How do I reset my password?",
    "How do I reset my password?",
    "how do i reset my password",
    "What's the way to reset my password?",
    "I forgot my password, what now?",
    "How many projects does the Free plan include?",
    "How many projects does the Pro plan include?",
    "How many projects are included in the Free plan?",
]


async def main():
    cache = SemanticCache()  # local embeddings, threshold 0.92, one hour TTL, per-user scope
    cache.warmup()
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={},
        instructions=INSTRUCTIONS,
        cache=cache,
    )
    alice = CallerContext(user_id="alice")
    total_tokens = 0
    try:
        for question in QUESTIONS:
            hits_before = (await cache.stats()).hits
            start = time.perf_counter()
            # Counts the tokens of every model call made inside this block.
            with get_usage_metadata_callback() as usage:
                await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
            elapsed_ms = (time.perf_counter() - start) * 1000
            hit = (await cache.stats()).hits > hits_before
            tokens = sum(u["total_tokens"] for u in usage.usage_metadata.values())
            total_tokens += tokens
            print(f"{'HIT ' if hit else 'MISS'} {elapsed_ms:7.0f} ms {tokens:6} tokens  {question}")

        stats = await cache.stats()
        print(f"\n{stats}  hit rate {stats.hit_rate:.0%}")
        print(f"Tokens used: {total_tokens}")
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
…
MISS    2990 ms    286 tokens  How do I reset my password?
HIT       17 ms      0 tokens  How do I reset my password?
HIT      620 ms      0 tokens  how do i reset my password
HIT       57 ms      0 tokens  What's the way to reset my password?
MISS    3185 ms    369 tokens  I forgot my password, what now?
MISS    1262 ms    136 tokens  How many projects does the Free plan include?
MISS    1854 ms    200 tokens  How many projects does the Pro plan include?
HIT       19 ms      0 tokens  How many projects are included in the Free plan?

CacheStats(hits=4, misses=4, stores=4, evictions=0)  hit rate 50%
Tokens used: 991

Half the questions cost nothing and came back in well under a second, most in under 60 ms, against one to three seconds for a model call. The Pro plan question missed even though it looks like the Free plan one, so nobody was told Pro has 3 projects. Your own hit rate depends entirely on how repetitive your traffic is. Measure it with cache.stats() on real questions before you count on savings.

Warning

In 1.2.1, don't combine cache= with observe=True. With observability on, every cache hit fails with TypeError: ObservabilityCollector.record() got an unexpected keyword argument 'description', the agent logs Cache check failed, continuing without cache, and calls the model anyway. The same eight questions with observe=True made 8 LLM calls and used 1,887 tokens, while cache.stats() still reported 4 hits. Count tokens with a callback, as above, until this is fixed.

Step 03

Keep each user's answers to themselves

A cached answer can contain anything the agent said to that user. That's why the default scope is per_user: each user_id gets its own partition, and a request without a caller isn't cached at all. Here's the part of isolation.py that checks it:

Pythonisolation.py
async def ask(agent, cache, who: str, caller: CallerContext | None) -> None:
    hits_before = (await cache.stats()).hits
    await agent.ainvoke(QUESTION, caller=caller)
    print(f"{'HIT ' if (await cache.stats()).hits > hits_before else 'MISS'}  {who}")
Pythonisolation.py
1
2
3
4
5
6
7
8
    alice, bob = CallerContext(user_id="alice"), CallerContext(user_id="bob")
    await ask(agent, cache, "alice", alice)
    await ask(agent, cache, "alice again", alice)
    await ask(agent, cache, "bob", bob)
    await ask(agent, cache, "no caller", None)
    await ask(agent, cache, "no caller again", None)
    print("purged entries for alice:", await cache.purge_user("alice"))
    await ask(agent, cache, "alice after purge", alice)
Output
scope='per_user' (the default)
…
MISS  alice
HIT   alice again
MISS  bob
MISS  no caller
MISS  no caller again
purged entries for alice: 1
MISS  alice after purge

Bob asked exactly what Alice asked and still got a fresh answer. purge_user removes everything cached for one user, which is what you call when someone asks to have their data erased. If your callers carry a tenant_id, it becomes part of the scope too, so two tenants with the same user ID never share entries; pass the same tenant_id to purge_user.

For finer isolation, scope="per_session" keys the cache on a session_id in the caller's metadata:

Pythonisolation.py
    cache = SemanticCache(scope="per_session")
    agent = await build_agent(model="openai:gpt-5-mini", servers={}, instructions=INSTRUCTIONS, cache=cache)
    morning = CallerContext(user_id="alice", metadata={"session_id": "chat-1"})
    evening = CallerContext(user_id="alice", metadata={"session_id": "chat-2"})
Output
scope='per_session'
…
MISS  alice, chat-1
MISS  alice, chat-2
HIT   alice, chat-1 again

The third scope, shared, gives every caller one cache and works without a CallerContext. Use it only when answers never depend on who's asking, such as a public FAQ, and pass shared_data_acknowledged=True, or Promptise logs a warning.

Step 04

Expire answers that go stale

Every entry has a time to live. default_ttl sets it in seconds, one hour by default. For questions about things that change, ttl_patterns maps a regular expression to a shorter TTL. Patterns are matched against the lowercased question, so write them in lowercase. This runs without a model, using the cache's own store and check:

Pythonttl_demo.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import asyncio

from promptise import CallerContext, SemanticCache


async def main():
    cache = SemanticCache(
        default_ttl=3600,  # most answers live for an hour
        ttl_patterns={r"status|today|price": 2},  # time-sensitive questions expire after 2 seconds
    )
    alice = CallerContext(user_id="alice")
    questions = ["How do I reset my password?", "Is there an outage today?", "What's the service status?"]
    for question in questions:
        await cache.store(question, "a stored answer", {"messages": []}, caller=alice)

    await asyncio.sleep(3)
    for question in questions:
        entry = await cache.check(question, caller=alice)
        print(f"{'HIT ' if entry else 'MISS'}  {question}")


asyncio.run(main())
Output
HIT   How do I reset my password?
MISS  Is there an outage today?
MISS  What's the service status?

In production you'd use something like 60 seconds rather than 2. The cache also misses on its own when you change the instructions or the model, because both are part of the key.


[05]

Survive LLM provider outages with model fallback, step by step

FallbackChain is a chat model that wraps several others. It tries them in order; if one raises an error or exceeds its timeout, the next one gets the same request. Each model has its own circuit breaker: after failure_threshold consecutive failures, the chain skips that model for recovery_timeout seconds, then lets one test request through.

Step 05

Simulate an outage you can switch on and off

You can't take OpenAI down on demand, so this small server stands in for it. While it's down it returns HTTP 503 to every request, which is what a provider outage looks like to your code. While it's up it forwards requests to the real OpenAI API:

Pythonoutage_proxy.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
"""A stand-in for your LLM provider that you can switch off and on.

While it's "down", every request gets HTTP 503, like a provider outage.
While it's "up", requests are forwarded to the real OpenAI API.
"""

import httpx
import uvicorn
from starlette.applications import Starlette
from starlette.requests import Request
from starlette.responses import JSONResponse, Response
from starlette.routing import Route

state = {"up": False, "requests": 0}


async def status(request: Request) -> JSONResponse:
    return JSONResponse(state)


async def switch(request: Request) -> JSONResponse:
    state["up"] = request.path_params["mode"] == "up"
    return JSONResponse(state)


async def provider(request: Request) -> Response:
    state["requests"] += 1
    if not state["up"]:
        return JSONResponse(
            {"error": {"message": "Service unavailable (simulated outage)", "type": "server_error"}},
            status_code=503,
        )
    async with httpx.AsyncClient(timeout=120) as client:
        upstream = await client.post(
            f"https://api.openai.com/v1/{request.path_params['path']}",
            content=await request.body(),
            headers={
                "authorization": request.headers["authorization"],
                "content-type": "application/json",
            },
        )
    return Response(upstream.content, upstream.status_code, media_type="application/json")


app = Starlette(
    routes=[
        Route("/admin/status", status, methods=["GET"]),
        Route("/admin/{mode:str}", switch, methods=["POST"]),
        Route("/v1/{path:path}", provider, methods=["POST"]),
    ]
)

if __name__ == "__main__":
    uvicorn.run(app, host="127.0.0.1", port=8190, log_level="warning")

Start it in its own terminal with python outage_proxy.py. Everything it needs comes with Promptise.

Note

Only an OpenAI key was available for this guide, so the backup model is another OpenAI model, gpt-4.1-mini. In production, put a different provider second, such as anthropic:..., so one provider's outage can't take out both. The chain treats every model the same way, but only OpenAI models were tested here.

Step 06

Put a fallback chain in front of the model

The primary is gpt-5-mini, pointed at the outage simulator with Model(..., endpoint=...). The backup talks to OpenAI directly. A low threshold and a short recovery time make the circuit breaker easy to watch:

Pythonfallback_demo.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
import asyncio
import time

import httpx

from promptise import FallbackChain, Model, build_agent

from acme_support import INSTRUCTIONS

PROXY = "http://127.0.0.1:8190"


def primary_requests() -> int:
    return httpx.get(f"{PROXY}/admin/status").json()["requests"]


async def main():
    chain = FallbackChain(
        [
            # The primary goes through the outage simulator. max_retries=0 stops the
            # OpenAI SDK from retrying on its own before the chain can fall back.
            Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
            "openai:gpt-4.1-mini",
        ],
        timeout_per_model=20,
        failure_threshold=2,
        recovery_timeout=10,
        on_fallback=lambda failed, backup, error: print(f"  fallback: {failed} -> {backup} ({error})"),
    )
    agent = await build_agent(model=chain, servers={}, instructions=INSTRUCTIONS)

    async def ask(question: str) -> None:
        before, start = primary_requests(), time.perf_counter()
        await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
        elapsed = time.perf_counter() - start
        state = chain.get_chain_status()[0]
        print(
            f"served by {chain.model_name:<20} {elapsed:5.2f}s  "
            f"primary tried {primary_requests() - before}x  primary circuit: {state['state']}"
        )

    try:
        httpx.post(f"{PROXY}/admin/down")
        print("Primary is down")
        await ask("How do I reset my password?")
        await ask("How many projects does the Free plan include?")
        await ask("When is human support available?")
        print(chain.get_chain_status())

        httpx.post(f"{PROXY}/admin/up")
        print("\nPrimary is back. Waiting for the recovery timeout...")
        await asyncio.sleep(10)
        await ask("How much does the Pro plan cost?")
        await ask("How much storage does the Pro plan include?")
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
…
Primary is down
  fallback: openai:gpt-5-mini -> openai:gpt-4.1-mini (Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}})
served by openai:gpt-4.1-mini   1.24s  primary tried 1x  primary circuit: closed
  fallback: openai:gpt-5-mini -> openai:gpt-4.1-mini (Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}})
served by openai:gpt-4.1-mini   0.56s  primary tried 1x  primary circuit: open
served by openai:gpt-4.1-mini   0.58s  primary tried 0x  primary circuit: open
[{'model_id': 'openai:gpt-5-mini', 'state': 'open', 'failures': 2, 'is_primary': True}, {'model_id': 'openai:gpt-4.1-mini', 'state': 'closed', 'failures': 0, 'is_primary': False}]

Primary is back. Waiting for the recovery timeout...
served by openai:gpt-5-mini     1.97s  primary tried 1x  primary circuit: closed
served by openai:gpt-5-mini     1.45s  primary tried 1x  primary circuit: closed

Read it from the top. The first two requests tried the primary, got a 503, and were answered by the backup. After the second failure the circuit opened, so the third request didn't touch the primary at all. Once the recovery timeout had passed, the next request was the test: the primary answered, the circuit closed, and traffic went back to it. on_fallback is the place to send an alert, and get_chain_status() is what a health check would report.

What counts as a failure? Any exception from the model, plus a timeout when you set timeout_per_model or global_timeout. That includes rate limits and server errors, but also errors that a second model won't fix, such as a request that's too long for the context window.

Step 07

Turn off the SDK's own retries

The OpenAI client retries failed requests by itself before it gives up, and the fallback can only start after that. retries_demo.py runs the same outage twice, once with the SDK's default and once with extra={"max_retries": 0} on the primary:

Pythonretries_demo.py
    for label, extra in [("SDK default retries", {}), ("max_retries=0", {"max_retries": 0})]:
        primary = Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra=extra)
        chain = FallbackChain([primary, "openai:gpt-4.1-mini"])
Output
…
SDK default retries   2.01s  primary tried 3x  served by openai:gpt-4.1-mini
…
max_retries=0         0.83s  primary tried 1x  served by openai:gpt-4.1-mini

With retries on, every request during an outage hits the dead provider three times before falling back. Turn retries off on every model except the last one in the chain, where a retry is your final chance.

Step 08

Let the cache answer when everything is down

The cache sits in front of the chain, so answers it already holds survive even a total outage. In outage_cache.py, both models go through the simulator. One question is asked while it's up, then everything goes down:

Pythonoutage_cache.py
1
2
3
4
5
6
7
8
9
    chain = FallbackChain(
        [
            Model("gpt-5-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
            Model("gpt-4.1-mini", provider="openai", endpoint=f"{PROXY}/v1", extra={"max_retries": 0}),
        ]
    )
    cache = SemanticCache()
    cache.warmup()
    agent = await build_agent(model=chain, servers={}, instructions=INSTRUCTIONS, cache=cache)
Output
…
Agent: Go to https://acme.example/reset, enter your account email, and follow the instructions sent to you. The reset link expires after 30 minutes, so use it promptly.

Every model is down
Agent: Go to https://acme.example/reset, enter your account email, and follow the instructions sent to you. The reset link expires after 30 minutes, so use it promptly.
GraphExecutionError: Graph 'react' failed at node 'reason': RuntimeError: All 2 models in FallbackChain failed.
  openai:gpt-5-mini: OpenAIAPIError: Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}}
  openai:gpt-4.1-mini: OpenAIAPIError: Error code: 503 - {'error': {'message': 'Service unavailable (simulated outage)', 'type': 'server_error'}}

The paraphrased password question still got an answer. The new question failed with an error that names every model and why it failed, which is what you want in your logs. Catch it and show the user a clear message.


[06]

Agents with tools: what changes

Most agents call tools, and both features behave differently there. This is where you need to be most careful.

The cache replays tool-using turns

Promptise caches a turn whether or not it called tools. On a hit, no tool runs; the agent returns the old answer. The tool names are saved with the entry, but nothing in 1.2.1 acts on them: invalidate_on_write=True is the default and the docs say a write tool evicts the cache, yet the agent never calls the method that does it. Here's a ticket agent with a read tool and a write tool, run with the default cache:

Pythontool_turns.py
1
2
3
4
5
6
7
8
9
10
        for question in [
            "How many open tickets do I have?",
            "Open a ticket: the invoice PDF is blank.",
            "How many open tickets do I have?",
            "Open a ticket: the invoice PDF is blank.",
        ]:
            print(f"\n>>> {question}")
            result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=alice)
            print("Agent:", result["messages"][-1].content)
            print("Tickets really open:", len(json.loads(DB.read_text())))
Output
>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Tickets really open: 2

>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-103", "subject": "Invoice PDF is blank"}
Agent: Ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3

>>> How many open tickets do I have?
Agent: You currently have 2 open support tickets.
Tickets really open: 3

>>> Open a ticket: the invoice PDF is blank.
Agent: Ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3

Two problems in four lines. The count is stale, because the write didn't clear the cache. Worse, the second "open a ticket" request never reached the tool, yet the agent said it did. For a tool that sends money or emails, that's a real bug in your product.

Until Promptise wires this up, clear the caller's cache yourself whenever a turn used a write tool. invalidate_for_write is public and removes every entry in that caller's scope, including the one just stored:

Pythontool_turns_fixed.py
1
2
3
4
5
6
7
8
9
10
11
WRITE_TOOLS = {"open_ticket"}
alice = CallerContext(user_id="alice")


async def ask(agent, cache: SemanticCache, question: str, caller: CallerContext) -> str:
    result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]}, caller=caller)
    tools_called = {call["name"] for message in result["messages"] for call in getattr(message, "tool_calls", None) or []}
    if tools_called & WRITE_TOOLS:
        # Something changed: drop this caller's cached answers, including the one just stored.
        await cache.invalidate_for_write(", ".join(sorted(tools_called & WRITE_TOOLS)), caller=caller)
    return result["messages"][-1].content
Output
>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Tickets really open: 2

>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-103", "subject": "Invoice PDF is blank"}
Agent: Done — ticket T-103 has been opened for "Invoice PDF is blank."
Tickets really open: 3

>>> How many open tickets do I have?
→ Invoking tool: count_open_tickets with {'payload': {}}
✔ Tool result from count_open_tickets: 3
Agent: You have 3 open support tickets.
Tickets really open: 3

>>> Open a ticket: the invoice PDF is blank.
→ Invoking tool: open_ticket with {'subject': 'Invoice PDF is blank'}
✔ Tool result from open_ticket: {"id": "T-104", "subject": "Invoice PDF is blank"}
Agent: I've opened ticket T-104 for "Invoice PDF is blank."
Tickets really open: 4

Both answers are right now. This only covers writes made through this agent; if data changes elsewhere, rely on a short TTL. The simplest rule is often the best one: put the cache on agents that answer questions, and leave it off agents that take actions.

FallbackChain doesn't work with tools in 1.2.1

To call tools, an agent binds them to the model. FallbackChain doesn't implement that yet, so any agent with tools fails on its first request:

Pythonfallback_tools.py
    agent = await build_agent(
        model=FallbackChain(["openai:gpt-5-mini", "openai:gpt-4.1-mini"]),
        servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
    )
Output
GraphExecutionError: Graph 'react' failed at node 'reason': NotImplementedError:

build_agent also accepts a LangChain runnable as the model, and LangChain's own with_fallbacks does support tool calling. In this test the primary's model name is deliberately wrong, so it fails every time:

Pythonfallback_tools_langchain.py
1
2
3
4
5
6
7
8
9
    model = ChatOpenAI(model="gpt-5-mini-does-not-exist", max_retries=0).with_fallbacks(
        [ChatOpenAI(model="gpt-4.1-mini", max_retries=0)]
    )
    agent = await build_agent(
        model=model,
        servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
        instructions="You are a support assistant. Use your tools. Answer in one sentence.",
        trace_tools=True,
    )
Output
→ Invoking tool: count_open_tickets with {'payload': None}
✔ Tool result from count_open_tickets: 2
Agent: You currently have 2 open support tickets.
Answered by: gpt-4.1-mini-2025-04-14

The backup answered and the tool call worked. What you give up is the circuit breaker: every request tries the primary first, so keep max_retries=0 and a timeout on it to make failures fast.


[07]

Honest limits

The cache saves real money on repeated questions, but a few cases need care. The first one is a bug in 1.2.1 that can give a user a wrong answer.

Multi-turn conversations can get the wrong cached answer. The cache embeds only the last message, and the context fingerprint counts the messages without looking at what they say. cache_multi_turn=False is the default, but 1.2.1 never reads it. So two conversations of the same length that end in the same follow-up collide, as multi_turn.py shows:

Output
…
Tell me about the city of Paris. / What river runs through it?
  Agent: The River Seine runs through Paris. 
  CacheStats(hits=0, misses=1, stores=1, evictions=0)
Tell me about the city of London. / What river runs through it?
  Agent: The River Seine runs through Paris. 
  CacheStats(hits=1, misses=1, stores=1, evictions=0)

The rest, in short:

  • Chat needs `scope="per_session"`. Then one conversation can't answer from another. Or cache only single-question requests.

  • Personalised answers are only safe per user. Never use scope="shared" for an agent that sees accounts, orders or anything private.

  • Time-sensitive answers need a short TTL. Prices, statuses and "today" questions belong in ttl_patterns, or out of the cache.

  • Tool calls with side effects are replayed, not run. Clear the cache after writes, as shown above, or keep the cache off agents that act.

  • The in-memory backend lives in one process. Each worker has its own cache, and a restart empties it. backend="redis" shares it across workers and survives restarts; it wasn't tested for this guide.

  • `CacheStats` can count one request twice. When a similar question is found but the context doesn't match, 1.2.1 counts both a hit and a miss. Treat the hit rate as approximate.


[08]

Frequently asked questions

What is semantic caching for LLMs?

A semantic cache stores model answers keyed by the meaning of the question rather than its exact text. Each question becomes an embedding, and a new question that's close enough to a cached one, by cosine similarity, gets the cached answer without a model call. That catches rephrasings and typos that an exact-match cache would miss.

How much does a semantic cache reduce LLM API costs?

Every hit saves a whole model call, so the saving equals your hit rate. In this guide's test, four of eight questions were hits and cost zero tokens. Real traffic varies a lot: FAQ-style support repeats itself, while open-ended chat rarely does. Run the cache on real questions and read cache.stats() before you estimate savings.

What similarity threshold should I use?

Start with the default 0.92 and test it on pairs from your own logs, as in Step 1. Raise it if you see wrong answers; lower it carefully if obvious rephrasings miss. Thresholds don't carry over between embedding models: the same lowercase question scored 0.994 with the local model and 0.908 with OpenAI's.

How does LLM fallback work with LangChain?

LangChain models have a with_fallbacks method that tries a list of backup runnables when the first one raises an error, and it supports tool calling. Promptise's FallbackChain adds per-model timeouts and a circuit breaker on top, but in 1.2.1 it only works for agents without tools. Both can be passed to build_agent(model=...).

Is a semantic cache safe when many users share one agent?

With the default per_user scope, yes: each user ID has its own partition, a request without a CallerContext isn't cached, and output guardrails run on every cached answer. The risks are stale or replayed answers, not leaks between users, and the limits above cover them.


[09]

Where to go next

  • Semantic Cache: every option, the Redis backend and custom embedding providers.

  • Model Fallback: FallbackChain options and circuit breaker states.

  • Models & Providers: Model and the provider strings for your backups.

  • Building Multi-User Systems: CallerContext, tenants and isolation across Promptise.

  • Events & Notifications: get alerted when requests fail.

  • How to Connect MCP Servers to Your AI Agent in Python: give the agent real tools.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths