A long agent run fills its context window one tool result at a time. Ten log files later, every model call carries all ten, you pay for them again on every step, and the agent's grip on the original question gets looser. Context engineering for AI agents is the work of deciding what the model sees on each step: what stays, what gets trimmed or summarized, and what it fetches only when it needs it. In this guide you'll build an agent that investigates an incident across ten service logs, meter the tokens of every model call, and compare Promptise Foundry's automatic compaction, its ContextEngine and a full transcript with real numbers, then fix what goes wrong. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed.
How do you manage the context window of an AI agent?
Keep three kinds of content apart. Rules and facts the agent needs on every step go in the system prompt. Bulk data stays behind tools that return only what the step needs, such as a summary instead of a whole file, with a second tool to drill into details. And the run itself needs a way to stay bounded as tool results pile up. With Promptise Foundry the first two are yours to write, and the third is built in:
agent = await build_agent(
model="openai:gpt-5-mini",
servers={"logs": StdioServerSpec(command=sys.executable, args=["lean_server.py"])},
instructions=INSTRUCTIONS,
)
meter = TokenMeter(watch="the number of ERROR lines per service")
try:
result = await agent.ainvoke(
# A HumanMessage, not a dict: in 1.2.1 it keeps the question in view after compaction.
{"messages": [HumanMessage(QUESTION)]},
config={"callbacks": [meter]},
)On the incident below, this setup read all ten logs and answered every part of the question correctly in three runs out of three, with 17,474 input tokens per run. The agent that read the raw files with nothing trimmed was just as correct, and used 323,477. The default setup, which compacts automatically, lost the question halfway through and never answered it. The rest of this guide shows each of those runs and why they behaved the way they did.
[02]
How context works in a long agent run
Every step of a tool loop is one model call. Without any management, each call sends the instructions, the question, and every tool call and result so far, so the input grows with each step. Promptise Foundry has two mechanisms that change what gets sent:
Rendering diagram…
Compaction inside the tool loop. The default agent pattern, ReAct, runs its loop with context_scope="auto". Until there are six tool results, the model sees the full transcript. From then on it sees a compact view: your instructions, the question, its most recent step, and a "facts already gathered" ledger with one line per tool call and its full result. agent_pattern="managed" uses that compact view from the first step. The Context Lifecycle guide and the Context scope reference describe both.
The ContextEngine, before the loop. build_agent(context_engine=...) assembles the instructions, chat history and memory into a token budget, and trims the least important parts first. It runs once per call to the agent, so it never sees the tool results produced during the run. The Context Engine docs list its layers.
What survives compaction matters more than how much it saves. This is what the model receives before and after the switch, from Promptise's source and the runs below:
In the prompt | Before compaction | After compaction, in 1.2.1 |
|---|---|---|
Your instructions | Yes | Yes |
The question, passed as a dict | Yes | No |
The question, passed as a HumanMessage | Yes | Yes, but only the first one in the input |
System messages you pass in | Yes | No |
Earlier turns of a chat | Yes | No, except the first user message |
Tool results | Every call and result | The last step's results, plus every result again in the ledger |
[03]
What you need
Python 3.10 or newer.
Promptise Foundry from PyPI. tiktoken, which the ContextEngine uses to count tokens exactly for OpenAI models, comes with it.
An OpenAI API key. This guide uses gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.
pip install promptise
export OPENAI_API_KEY="sk-..."[04]
Investigate an incident across ten logs, step by step
The job: an incident hit checkout between 14:00 and 14:30. The agent reads the logs of ten services and reports the ERROR count per service, the total, the first ERROR, and the root cause. A small script, make_logs.py, wrote the logs (140 normal lines per service plus the incident) and printed the right answer, so every run can be checked against it:
files: 10 | bytes: 135617
ERROR lines per service: {'gateway': 4, 'orders': 5, 'payments': 2, 'inventory': 7}
total ERROR lines: 18
first ERROR: inventory 14:07:41.204
…Give the agent the logs as a tool
The tools live in a small MCP server. read_log returns a whole file, the way a first version of a tool usually does:
from pathlib import Path
from promptise.mcp.server import MCPServer
server = MCPServer("logs")
LOG_DIR = Path(__file__).parent / "logs"
def services() -> list[str]:
return sorted(p.stem for p in LOG_DIR.glob("*.log"))
@server.tool()
async def list_logs() -> list[str]:
"""List the services that have a log file for the incident window."""
return services()
@server.tool()
async def read_log(service: str) -> str:
"""Return the full log file of one service, every line.
Args:
service: A service name from list_logs, for example "orders".
"""
if service not in services():
return f"No log for service {service!r}. Call list_logs for valid names."
return (LOG_DIR / f"{service}.log").read_text()
if __name__ == "__main__":
server.run()How big is one result? The ContextEngine's token counter answers that without calling a model:
"""How big is each tool result? Count before you send it."""
from pathlib import Path
from promptise import ContextEngine
engine = ContextEngine(model="openai:gpt-5-mini")
for path in sorted(Path("logs").glob("*.log"))[:3]:
print(f"{path.name:<16} {engine.count_tokens(path.read_text()):>6,} tokens")auth.log 5,557 tokens
billing.log 5,848 tokens
gateway.log 5,829 tokensAbout 5,800 tokens per file, 58,000 for all ten. If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python explains this file line by line.
Meter every model call
You can't manage what you don't measure. This LangChain callback records, for each model call, how many messages went in, whether the compaction ledger was among them, whether a phrase you choose was still visible, and the token usage the provider reported:
"""Record what the model saw and what it cost, one row per model call."""
from langchain_core.callbacks import AsyncCallbackHandler
LEDGER_MARKER = "Facts already gathered"
class TokenMeter(AsyncCallbackHandler):
def __init__(self, watch: str | None = None):
self.watch = watch # a phrase to look for in what the model sees
self.calls: list[dict] = []
async def on_chat_model_start(self, serialized, messages, **kwargs):
sent = messages[0]
text = "\n".join(str(m.content) for m in sent)
self.calls.append(
{
"messages": len(sent),
"tool_results": sum(1 for m in sent if m.type == "tool"),
"ledger": LEDGER_MARKER in text,
"watch_seen": bool(self.watch) and self.watch in text,
}
)
async def on_llm_end(self, response, **kwargs):
message = response.generations[0][0].message
usage = message.usage_metadata or {}
self.calls[-1].update(
input=usage.get("input_tokens", 0),
cached=(usage.get("input_token_details") or {}).get("cache_read", 0),
output=usage.get("output_tokens", 0),
wants=", ".join(
tc["name"] + "".join(f" {v}" for v in tc["args"].values()) for tc in message.tool_calls
)
or "final answer",
)The file also has a table() method that prints one row per call and a total, priced at OpenAI's list price for gpt-5-mini: $0.25 per million input tokens, $0.025 per million cached input tokens and $2.00 per million output tokens. You pass the meter to any run with config={"callbacks": [meter]}. For tracing in production, Promptise's own observability records tokens per call too; the meter is here because it also shows what the model was sent.
Run the long investigation with the default agent
The script asks the question, watches for the phrase "the number of ERROR lines per service" from it, and can build the agent four ways. The instructions ask for one service per step, so the run is long enough to watch it grow:
"""Investigate an incident across ten service logs, and meter every model call.
python investigate.py # the default agent
python investigate.py managed # agent_pattern="managed"
python investigate.py full # no compaction: the whole transcript every call
python investigate.py engine # the default agent with a ContextEngine
python investigate.py default lean_server.py # the same agent with smaller tool results
"""
import asyncio
import sys
from contextlib import AsyncExitStack
from promptise import ContextEngine, build_agent
from promptise.config import StdioServerSpec
from promptise.engine import PromptGraph, PromptNode
from promptise.mcp.client import MCPClient, MCPMultiClient, MCPToolAdapter
from meter import TokenMeter
INSTRUCTIONS = (
"You are an SRE assistant. Answer from the logs only. "
"Work through the logs one service at a time: call one tool for one service, "
"look at the result, then go to the next. Never ask about two services in one step."
)
QUESTION = (
"We had an incident on 2026-10-09 between 14:00 and 14:30. Read every service log. "
"Then tell me: (1) the number of ERROR lines per service, (2) the total, "
"(3) which service logged the first ERROR and its exact timestamp, "
"and (4) the root cause in one sentence."
)
SERVER_FILE = sys.argv[2] if len(sys.argv) > 2 else "logs_server.py"
SERVER = StdioServerSpec(command=sys.executable, args=[SERVER_FILE])
async def full_transcript_graph(stack: AsyncExitStack) -> PromptGraph:
"""A ReAct graph whose node always sees the whole transcript."""
client = MCPMultiClient({"logs": MCPClient(transport="stdio", command=sys.executable, args=[SERVER_FILE])})
await stack.enter_async_context(client)
tools = await MCPToolAdapter(client).as_langchain_tools()
graph = PromptGraph("full-transcript", mode="static")
graph.add_node(
PromptNode(
"reason",
instructions=INSTRUCTIONS,
tools=tools,
context_scope="full",
max_iterations=15,
default_next="__end__",
)
)
graph.set_entry("reason")
return graph
async def main(mode: str) -> None:
async with AsyncExitStack() as stack:
if mode == "full":
options = {"servers": {}, "agent_pattern": await full_transcript_graph(stack)}
elif mode == "engine":
options = {"servers": {"logs": SERVER}, "context_engine": ContextEngine(model="openai:gpt-5-mini")}
elif mode == "managed":
options = {"servers": {"logs": SERVER}, "agent_pattern": "managed"}
else:
options = {"servers": {"logs": SERVER}}
agent = await build_agent(model="openai:gpt-5-mini", instructions=INSTRUCTIONS, **options)
meter = TokenMeter(watch="the number of ERROR lines per service")
try:
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": QUESTION}]},
config={"callbacks": [meter]},
)
print(meter.table())
print("\n>>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main(sys.argv[1] if len(sys.argv) > 1 else "default"))python investigate.pycall msgs tool msgs ledger input cached output next
1 2 0 no 357 0 87 list_logs {} [watch: seen]
2 4 1 no 415 0 87 read_log auth [watch: seen]
3 6 2 no 6,000 5,888 87 read_log billing [watch: seen]
4 8 3 no 11,876 11,648 17 read_log gateway [watch: seen]
5 10 4 no 17,733 17,536 17 read_log inventory [watch: seen]
6 12 5 no 23,702 23,040 17 read_log notifications [watch: seen]
7 4 1 yes 34,313 0 615 final answer [watch: MISSING]
total: 7 model calls, 94,396 input (58,112 cached) + 927 output tokens = $0.0124
>>> Which service should I investigate first? Available logs: auth, billing, gateway, inventory, notifications, orders, payments, search, shipping, users.Read it row by row. Each log added about 5,900 input tokens to every later call. After six tool results, list_logs and five logs, compaction switched on at call 7. The message count dropped from 12 to 4: the instructions, the last tool call, its result, and the ledger. The question was not among them. The model saw an SRE assistant's instructions and a pile of logs, and asked what to do. It never read the last five services.
The cause is in the compaction code: it keeps the first message in the input that is a LangChain HumanMessage object. A plain {"role": "user", ...} dict, the form used in the Promptise docs, isn't one, so the question is dropped. Notice also that call 7 is bigger than it would be without compaction: 34,313 tokens, where the full transcript in the next step reaches 29,088 at the same point. The latest result is sent twice, once in the step and once in the ledger.
Compare it with the full transcript
To see what compaction should be saving, run the same agent with compaction turned off. There's no build_agent switch for that in 1.2.1, so full_transcript_graph above builds the ReAct graph by hand with context_scope="full", using tools loaded through Promptise's MCP client:
python investigate.py full[promptise] No tools discovered from MCP servers; agent will run without tools.
call msgs tool msgs ledger input cached output next
1 2 0 no 357 0 87 list_logs {} [watch: seen]
2 4 1 no 415 0 87 read_log auth [watch: seen]
3 6 2 no 6,000 5,888 151 read_log billing [watch: seen]
…
10 20 9 no 46,909 46,592 17 read_log shipping [watch: seen]
11 22 10 no 52,579 51,712 17 read_log users [watch: seen]
12 24 11 no 58,622 58,496 1,855 final answer [watch: seen]
total: 12 model calls, 323,477 input (318,080 cached) + 2,316 output tokens = $0.0139
>>> Summary from logs (2026-10-09 14:00–14:30):
1) ERROR lines per service:
- auth: 0
- billing: 0
- gateway: 4
- inventory: 7
…
2) Total ERROR lines: 18
3) First ERROR: inventory at 2026-10-09T14:07:41.204Z
4) Root cause (one sentence): Inventory service exhausted its database connection pool (logs show "db pool exhausted (max=20, waiting=...)"), causing inventory HTTP 500s that propagated to orders ("inventory reserve/stock lookup failed: HTTP 500") and then to the gateway (502 upstream errors).The "No tools discovered" line is expected: the tools come with the graph, not through servers. Every log was read and every part of the answer is right. The price is growth: the last call sent 58,622 tokens, and the run sent 323,477 in total, because every result rides along on every later call.
One thing kept the bill low here. The cached column shows input the provider served from its prompt cache: a full transcript only ever appends, so each call starts with the previous call's exact text. At list price without the cache, the same run costs about $0.085. Cache hits vary from run to run, and they don't help the context window: with a hundred logs instead of ten, this design simply stops fitting.
Make the tools return less
The cheapest token is the one a tool never returns. The agent doesn't need 140 lines of 200 OK to count errors; it needs the counts and the error lines. So summarize where the data lives, and give the model a second tool to drill in when it needs detail:
from collections import Counter
from pathlib import Path
from promptise.mcp.server import MCPServer
server = MCPServer("logs")
LOG_DIR = Path(__file__).parent / "logs"
def services() -> list[str]:
return sorted(p.stem for p in LOG_DIR.glob("*.log"))
def lines_of(service: str) -> list[str]:
return (LOG_DIR / f"{service}.log").read_text().splitlines()
@server.tool()
async def list_logs() -> list[str]:
"""List the services that have a log file for the incident window."""
return services()
@server.tool()
async def log_summary(service: str) -> dict:
"""Summarize one service's log: line counts per level, time range, and every ERROR line.
Args:
service: A service name from list_logs, for example "orders".
"""
if service not in services():
return {"error": f"No log for service {service!r}. Call list_logs for valid names."}
lines = lines_of(service)
return {
"service": service,
"lines": len(lines),
"levels": dict(Counter(line.split()[1] for line in lines)),
"from": lines[0].split()[0],
"to": lines[-1].split()[0],
"errors": [line for line in lines if line.split()[1] == "ERROR"],
}
@server.tool()
async def search_log(service: str, text: str, limit: int = 20) -> list[str]:
"""Return up to `limit` lines of one service's log that contain `text`, such as a request ID.
Args:
service: A service name from list_logs.
text: Text to look for, for example "r-4be19a" or "WARN".
limit: The most lines to return.
"""
if service not in services():
return [f"No log for service {service!r}. Call list_logs for valid names."]
return [line for line in lines_of(service) if text in line][:limit]
if __name__ == "__main__":
server.run()The counting now happens in code, which counts correctly every time, and the limit on search_log keeps a drill-down from flooding the context. Same default agent, smaller tools:
python investigate.py default lean_server.pycall msgs tool msgs ledger input cached output next
1 2 0 no 500 0 151 list_logs {} [watch: seen]
2 4 1 no 558 0 87 log_summary auth [watch: seen]
…
6 12 5 no 1,432 1,280 17 log_summary notifications [watch: seen]
7 4 1 yes 1,464 0 407 log_summary orders [watch: MISSING]
8 4 1 yes 1,954 1,792 471 log_summary payments [watch: MISSING]
…
11 4 1 yes 2,078 0 343 log_summary users [watch: MISSING]
12 4 1 yes 2,160 0 1,891 final answer [watch: MISSING]
total: 12 model calls, 16,574 input (6,656 cached) + 3,714 output tokens = $0.0101
>>> Summary
- Primary failure began ~2026-10-09T14:07:41 and caused checkout/orders errors for clients.
- Root cause visible in logs: inventory service returned HTTP 500 due to "db pool exhausted (max=20, waiting=...)" which caused orders' inventory calls to fail and the gateway to return 502s.
…
If you want, I can fetch specific log lines around any request ID (for example r-4be19a or r-c27d18) or pull more inventory log context to help identify why connections were held. Which request ID or service should I look at next?The input fell from 323,477 to 16,574 tokens, almost twenty times less, and the agent got through all ten services. But the question still disappeared at call 7, so the answer is a generic incident summary with no counts and no total. Smaller results fix the cost; they don't fix what compaction drops.
Keep the question and the facts that must survive
Two changes keep what matters in view. Pass the question as a HumanMessage, which the compaction step keeps. And put the rules and facts the agent needs at its last step into instructions, because they are the only part of your input that is always sent. Here that's who the report is for and what shape it takes:
"""The same long run, set up to keep what matters."""
import asyncio
import sys
from langchain_core.messages import HumanMessage
from promptise import build_agent
from promptise.config import StdioServerSpec
from meter import TokenMeter
# Rules and facts the agent needs on every step live in the instructions.
INSTRUCTIONS = """\
You are an SRE assistant. Answer from the logs only.
Work through the logs one service at a time: call one tool for one service,
look at the result, then go to the next. Never ask about two services in one step.
Facts for every report:
- Incident commander: Priya Raman. Address the report to her by name.
- Answer exactly the numbered parts the user asks for. No mitigation advice.
"""
QUESTION = (
"We had an incident on 2026-10-09 between 14:00 and 14:30. Read every service log. "
"Then tell me: (1) the number of ERROR lines per service, (2) the total, "
"(3) which service logged the first ERROR and its exact timestamp, "
"and (4) the root cause in one sentence."
)
async def main() -> None:
agent = await build_agent(
model="openai:gpt-5-mini",
servers={"logs": StdioServerSpec(command=sys.executable, args=["lean_server.py"])},
instructions=INSTRUCTIONS,
)
meter = TokenMeter(watch="the number of ERROR lines per service")
try:
result = await agent.ainvoke(
# A HumanMessage, not a dict: in 1.2.1 it keeps the question in view after compaction.
{"messages": [HumanMessage(QUESTION)]},
config={"callbacks": [meter]},
)
print(meter.table())
print("\n>>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())python final.pycall msgs tool msgs ledger input cached output next
1 2 0 no 536 0 87 list_logs {} [watch: seen]
2 4 1 no 594 0 87 log_summary auth [watch: seen]
3 6 2 no 694 0 17 log_summary billing [watch: seen]
4 8 3 no 794 0 17 log_summary gateway [watch: seen]
5 10 4 no 1,053 0 17 log_summary inventory [watch: seen]
6 12 5 no 1,468 0 17 log_summary notifications [watch: seen]
7 5 1 yes 1,578 0 151 log_summary orders [watch: seen]
8 5 1 yes 2,068 1,920 215 log_summary payments [watch: seen]
9 5 1 yes 2,113 0 151 log_summary search [watch: seen]
10 5 1 yes 2,110 0 215 log_summary shipping [watch: seen]
11 5 1 yes 2,192 0 791 log_summary users [watch: seen]
12 5 1 yes 2,274 0 985 final answer [watch: seen]
total: 12 model calls, 17,474 input (1,920 cached) + 2,750 output tokens = $0.0094
>>> Priya Raman,
1) Number of ERROR lines per service:
- auth: 0
- billing: 0
- gateway: 4
- inventory: 7
- notifications: 0
- orders: 5
- payments: 2
- search: 0
- shipping: 0
- users: 0
2) Total: 18
3) First ERROR: inventory at 2026-10-09T14:07:41.204Z
4) Root cause (one sentence): The inventory service exhausted its database connection pool (db pool exhausted), returning HTTP 500s that caused order reservation failures and cascading 502 errors at the gateway.After compaction the model now gets five messages, and the question is one of them on every call. The answer matches the ground truth line by line, the report is addressed to the incident commander from the instructions, and the whole run cost under a cent. Each call stayed between 536 and 2,274 input tokens, so a run over a hundred services would still fit comfortably.
[05]
The numbers from every run
Each setup ran three times on the same logs and question. "Complete" means all four parts of the answer matched the ground truth.
Setup | Model calls | Input tokens | Logs read | Complete |
|---|---|---|---|---|
Default agent, whole files | 7 to 10 | 94,396 to 234,737 | 5 to 8 | 0 of 3 |
agent_pattern="managed", whole files | 5 to 7 | 55,641 to 120,681 | 3 to 5 | 0 of 3 |
Default agent with a ContextEngine, whole files | 8 to 9 | 135,914 to 182,421 | 6 to 7 | 0 of 3 |
No compaction, whole files | 12 | 323,477 | 10 | 3 of 3 |
Default agent, summaries | 12 | 16,574 to 17,074 | 9 to 10 | 0 of 3 |
Default agent with a ContextEngine, summaries | 9 to 12 | 10,664 to 16,862 | 7 to 9 | 0 of 3 |
No compaction, summaries | 12 | 16,971 | 10 | 3 of 3 |
final.py: summaries, HumanMessage, facts in instructions | 12 | 17,474 | 10 | 3 of 3 |
The compacted runs that used fewer tokens did so by stopping early, not by compressing. Every setup that kept the question in view got the answer right, and the tool design decided the cost.
[06]
When the model reads all ten logs at once
Models often ask for several tools in one step. With the instruction changed to "Read each service's log with read_log, one service per call." in investigate_parallel.py, gpt-5-mini asked for all ten logs in its second call. Here's the default agent, then the full transcript:
call msgs tool msgs ledger input cached output next
1 2 0 no 334 0 23 list_logs {} [watch: seen]
2 4 1 no 392 0 229 read_log auth, read_log billing, read_lo [watch: seen]
3 13 10 yes 116,479 116,352 2,261 final answer [watch: MISSING]
total: 3 model calls, 117,205 input (116,352 cached) + 2,513 output tokens = $0.0081call msgs tool msgs ledger input cached output next
1 2 0 no 334 0 23 list_logs {} [watch: seen]
2 4 1 no 392 0 229 read_log auth, read_log billing, read_lo [watch: seen]
3 15 11 no 58,518 0 1,578 final answer [watch: seen]
total: 3 model calls, 59,244 input (0 cached) + 1,830 output tokens = $0.0185With eleven results in one step, the compacted call carried all ten logs twice, once as the last step and once in the ledger: 116,479 tokens against 58,518. And again the question was gone. Batching cuts the number of calls, but it puts every result into one call, so the size of each result matters even more.
[07]
Budget the starting context with the ContextEngine
The ContextEngine solves a different problem: a chat whose history keeps growing. Before each call it counts the instructions, history, memory and question with tiktoken, and when they exceed the budget it drops the least important parts first, oldest conversation turns before anything else. This example uses a tiny window so trimming shows on a short history:
"""Budget the context an agent starts with: instructions, chat history and the new question."""
import asyncio
from langchain_core.callbacks import AsyncCallbackHandler
from promptise import ContextEngine, build_agent
class ShowPrompt(AsyncCallbackHandler):
"""Print every message the model receives on its first call."""
def __init__(self):
self.done = False
async def on_chat_model_start(self, serialized, messages, **kwargs):
if not self.done:
self.done = True
for m in messages[0]:
print(f" {m.type:<6} {str(m.content)[:60]}")
# Six earlier turns of a chat about the incident.
history = []
for service in ["auth", "billing", "gateway", "inventory", "notifications", "orders"]:
history.append({"role": "user", "content": f"Summarize the {service} log."})
history.append(
{"role": "assistant", "content": f"The {service} log has 140 INFO lines. " + "Nothing unusual. " * 40}
)
async def main() -> None:
default = ContextEngine(model="openai:gpt-5-mini")
print(f"default for gpt-5-mini: window {default.window:,}, budget {default.budget:,}")
# A deliberately tiny window, so trimming happens on a short example.
engine = ContextEngine(model="openai:gpt-5-mini", model_context_window=800, response_reserve=300)
engine.add_layer("runbook", priority=7, required=True, content="Incident commander: Priya Raman.")
agent = await build_agent(
model="openai:gpt-5-mini",
servers={},
instructions="You are an SRE assistant. Be brief.",
context_engine=engine,
)
try:
print("the model sees:")
result = await agent.ainvoke(
{"messages": history + [{"role": "user", "content": "Which services have we covered so far?"}]},
config={"callbacks": [ShowPrompt()]},
)
report = engine.last_report
print(f"budget {report.budget}, used {report.total_tokens} ({report.utilization:.0%}), trimmed {report.trimmed_layers}")
for layer in report.layers:
print(f" {layer['name']:<13} priority {layer['priority']:>2} {layer['tokens']:>4} tokens trimmed={layer['trimmed']}")
print(">>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())default for gpt-5-mini: window 128,000, budget 123,904
[promptise] No tools discovered from MCP servers; agent will run without tools.
the model sees:
system You are an SRE assistant. Be brief.
system You are an SRE assistant. Be brief.
human Summarize the inventory log.
ai The inventory log has 140 INFO lines. Nothing unusual. Nothi
human Summarize the notifications log.
ai The notifications log has 140 INFO lines. Nothing unusual. N
human Summarize the orders log.
ai The orders log has 140 INFO lines. Nothing unusual. Nothing
human Which services have we covered so far?
budget 500, used 438 (88%), trimmed ['conversation']
identity priority 10 10 tokens trimmed=False
conversation priority 1 420 tokens trimmed=True
user_message priority 10 8 tokens trimmed=FalseWhat worked: the budget is the window minus response_reserve, the three oldest exchanges were dropped as whole user and assistant pairs, the new question stayed, and last_report says exactly what happened. The agent now believes only three services were covered, which is what trimming means: dropped turns are gone, not summarized. What didn't: the instructions were sent twice, and the runbook layer, marked required=True, never reached the model. More on both in the limits below.
For chats stored through a conversation store, there's a simpler cap: build_agent(conversation_max_messages=N) keeps only the newest N messages of each session when it saves them.
[08]
Practical rules for agent context
These held up across every run in this guide:
Measure before you tune. Count a tool result with ContextEngine.count_tokens and meter real runs. The numbers decide where the work is; here it was entirely in the tool results.
Return less from tools. Summarize, filter and count where the data lives, in code. Return the error lines, not the file. Add a drill-down tool with a limit for the cases where the model needs detail.
System prompt for what must always hold. Output format, who the report is for, hard rules, the facts the last step depends on. It's sent on every call, twelve times in this run, so keep it short. In 1.2.1 it's also the only input that survives compaction.
Tools for everything large or occasional. Logs, runbooks, past incidents and documentation belong behind a search or lookup tool, the way RAG turns documents into a search tool. The model fetches the slice it needs, when it needs it.
Summaries are yours to write. Promptise 1.2.1 doesn't ask a model to summarize the run: the ledger repeats every tool result word for word, and the ContextEngine drops whole turns. If a summary matters, produce it in the tool, as log_summary does.
Keep one long job per question. Compaction keeps the first question in the input, so run long investigations as their own ainvoke call with a single HumanMessage, and bring the result back into the chat afterwards.
[09]
Honest limits
These are true of Promptise Foundry 1.2.1, and worth knowing before a long run goes to production:
Compaction drops a question passed as a dict. After six tool results, the default agent and agent_pattern="managed" look for a HumanMessage object to keep, and a {"role": "user"} dict isn't one. Pass HumanMessage objects, as final.py does.
Compaction keeps the first question, not the latest. With chat history in the input, as agent.chat() builds it, a compacted turn sees the first user message of the session. In a test with one earlier exchange, the agent answered the earlier question ("Which services have logs?") instead of the incident question. Keep tool loops inside a chat turn under six results, or run long jobs as separate calls.
Compaction drops system messages you pass in. Anything injected as a system message, such as a runtime's context state, disappears after six tool results. Facts the run needs go in instructions.
The ledger doesn't shrink unique results. It removes duplicate calls, but every distinct result appears in full, and the latest step's results appear twice. On a batch of ten logs that doubled the final call.
There's no switch to turn compaction off. build_agent has no context_scope option. Building the graph yourself works, as investigate.py full shows, with two catches: a single-node PromptGraph passed to agent_pattern needs mode="static", or Promptise wraps it as an autonomous graph that keeps looping after the answer, and a node with inject_tools=True receives no MCP tools this way, so load the tools with MCPToolAdapter and pass them to the node.
ContextEngine custom layers don't reach the model. The agent clears every layer before each call and refills only the built-in ones, so content from add_layer(...) is lost, required=True included. The instructions are also sent twice, once as the identity layer and once by the agent, and the engine's question is a plain dict, so a ContextEngine agent loses its question at compaction even when you pass a HumanMessage. Use it for chats with long histories and short tool loops.
The ContextEngine never sees tool results. It runs once before the tool loop. Its tools layer isn't filled either, so tool schemas aren't counted against the budget.
The default window for gpt-5-mini is 128,000 tokens. OpenAI lists 400,000 for gpt-5-mini. Pass model_context_window when you want the engine to budget for the real window.
Two docs pages are behind the code. The Context Lifecycle guide says the default ReAct pattern uses context_scope="full"; the 1.2.1 code uses "auto", as the Context scope reference says. The Context Engine page shows register_layer(...), which doesn't exist in 1.2.1; the method is add_layer(...).
[10]
Frequently asked questions
What is context engineering for AI agents?
Deciding what goes into the model's context on each step of an agent run: instructions, the task, history, tool results and retrieved data, and in what form. Prompt engineering is about the wording of one prompt; context engineering is about the whole working set as it changes over a long run. In practice most of it is tool design: what each tool returns, and how much.
What is context compaction?
Replacing a long transcript with a shorter view before the next model call. Promptise's default agent does it automatically after six tool results, by keeping the instructions, the first question, the latest step and a ledger of every tool result. Other systems summarize the history with a model instead. Either way, check what the compacted view still contains, because that's all the model knows.
Does a bigger context window solve the problem?
Not on its own. The full-transcript run here peaked at 58,622 tokens, well inside gpt-5-mini's window, and answered correctly, but it sent 323,477 input tokens to do work that took 17,474 with smaller tool results. Every token in the window is paid for on every call, and the growth doesn't stop at ten files.
How do I reduce the token usage of an AI agent?
Meter each model call first, then shrink the largest input. For agents that's almost always tool results: return summaries and counts instead of raw data, cap list sizes, and add a drill-down tool. Keep the system prompt short, since it's resent every call. Prompt caching lowers the price of repeated prefixes, but not the size of the context.
Should I summarize conversation history?
When a chat outgrows its budget, something has to give: either old turns are dropped, as the ContextEngine and conversation_max_messages do, or they're summarized. Promptise 1.2.1 doesn't summarize history for you. If older turns matter, keep the facts they produced in a tool or in the instructions, rather than relying on the turns themselves.
[11]
Where to go next
Context Engine: layers, priorities, budgets and the assembly report.
Context Lifecycle Management: the full, scoped and ledger modes and the managed pattern.
Engine nodes: Context scope: every PromptNode option, including context_scope and auto_ledger_after.
Reasoning Patterns: the built-in agent patterns and when each one fits.
Context System for prompts: context providers that inject user, environment and error context into prompts.
Observability: record tokens and latency for every call in production.
How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.
OpenAPI to MCP: Turn Any REST API into an MCP Server: generated tools are a common source of oversized results, and a good place to apply these rules.