PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Context Engineering for AI Agents: Context Window Management

Context engineering for AI agents in Python: keep what matters in long runs, trim the rest, and cut input tokens 18x. Real token counts per step.

Level
Intermediate
Reading time
26 min
Published
Oct 10, 2026
By
Promptise Team
  • Context Engineering
  • Context Window
  • AI Agents
  • Token Usage
  • MCP
  • Promptise Foundry

A long agent run fills its context window one tool result at a time. Ten log files later, every model call carries all ten, you pay for them again on every step, and the agent's grip on the original question gets looser. Context engineering for AI agents is the work of deciding what the model sees on each step: what stays, what gets trimmed or summarized, and what it fetches only when it needs it. In this guide you'll build an agent that investigates an incident across ten service logs, meter the tokens of every model call, and compare Promptise Foundry's automatic compaction, its ContextEngine and a full transcript with real numbers, then fix what goes wrong. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed.

[01]

How do you manage the context window of an AI agent?

Keep three kinds of content apart. Rules and facts the agent needs on every step go in the system prompt. Bulk data stays behind tools that return only what the step needs, such as a summary instead of a whole file, with a second tool to drill into details. And the run itself needs a way to stay bounded as tool results pile up. With Promptise Foundry the first two are yours to write, and the third is built in:

Pythonfinal.py
1
2
3
4
5
6
7
8
9
10
11
12
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"logs": StdioServerSpec(command=sys.executable, args=["lean_server.py"])},
        instructions=INSTRUCTIONS,
    )
    meter = TokenMeter(watch="the number of ERROR lines per service")
    try:
        result = await agent.ainvoke(
            # A HumanMessage, not a dict: in 1.2.1 it keeps the question in view after compaction.
            {"messages": [HumanMessage(QUESTION)]},
            config={"callbacks": [meter]},
        )

On the incident below, this setup read all ten logs and answered every part of the question correctly in three runs out of three, with 17,474 input tokens per run. The agent that read the raw files with nothing trimmed was just as correct, and used 323,477. The default setup, which compacts automatically, lost the question halfway through and never answered it. The rest of this guide shows each of those runs and why they behaved the way they did.


[02]

How context works in a long agent run

Every step of a tool loop is one model call. Without any management, each call sends the instructions, the question, and every tool call and result so far, so the input grows with each step. Promptise Foundry has two mechanisms that change what gets sent:

Rendering diagram…

  • Compaction inside the tool loop. The default agent pattern, ReAct, runs its loop with context_scope="auto". Until there are six tool results, the model sees the full transcript. From then on it sees a compact view: your instructions, the question, its most recent step, and a "facts already gathered" ledger with one line per tool call and its full result. agent_pattern="managed" uses that compact view from the first step. The Context Lifecycle guide and the Context scope reference describe both.

  • The ContextEngine, before the loop. build_agent(context_engine=...) assembles the instructions, chat history and memory into a token budget, and trims the least important parts first. It runs once per call to the agent, so it never sees the tool results produced during the run. The Context Engine docs list its layers.

What survives compaction matters more than how much it saves. This is what the model receives before and after the switch, from Promptise's source and the runs below:

In the prompt

Before compaction

After compaction, in 1.2.1

Your instructions

Yes

Yes

The question, passed as a dict

Yes

No

The question, passed as a HumanMessage

Yes

Yes, but only the first one in the input

System messages you pass in

Yes

No

Earlier turns of a chat

Yes

No, except the first user message

Tool results

Every call and result

The last step's results, plus every result again in the ledger


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry from PyPI. tiktoken, which the ContextEngine uses to count tokens exactly for OpenAI models, comes with it.

  • An OpenAI API key. This guide uses gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.

>_Terminal
pip install promptise
export OPENAI_API_KEY="sk-..."

[04]

Investigate an incident across ten logs, step by step

The job: an incident hit checkout between 14:00 and 14:30. The agent reads the logs of ten services and reports the ERROR count per service, the total, the first ERROR, and the root cause. A small script, make_logs.py, wrote the logs (140 normal lines per service plus the incident) and printed the right answer, so every run can be checked against it:

Output
files: 10 | bytes: 135617
ERROR lines per service: {'gateway': 4, 'orders': 5, 'payments': 2, 'inventory': 7}
total ERROR lines: 18
first ERROR: inventory 14:07:41.204
…
Step 01

Give the agent the logs as a tool

The tools live in a small MCP server. read_log returns a whole file, the way a first version of a tool usually does:

Pythonlogs_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
from pathlib import Path

from promptise.mcp.server import MCPServer

server = MCPServer("logs")
LOG_DIR = Path(__file__).parent / "logs"


def services() -> list[str]:
    return sorted(p.stem for p in LOG_DIR.glob("*.log"))


@server.tool()
async def list_logs() -> list[str]:
    """List the services that have a log file for the incident window."""
    return services()


@server.tool()
async def read_log(service: str) -> str:
    """Return the full log file of one service, every line.

    Args:
        service: A service name from list_logs, for example "orders".
    """
    if service not in services():
        return f"No log for service {service!r}. Call list_logs for valid names."
    return (LOG_DIR / f"{service}.log").read_text()


if __name__ == "__main__":
    server.run()

How big is one result? The ContextEngine's token counter answers that without calling a model:

Pythoncount_tokens.py
1
2
3
4
5
6
7
8
"""How big is each tool result? Count before you send it."""
from pathlib import Path

from promptise import ContextEngine

engine = ContextEngine(model="openai:gpt-5-mini")
for path in sorted(Path("logs").glob("*.log"))[:3]:
    print(f"{path.name:<16} {engine.count_tokens(path.read_text()):>6,} tokens")
Output
auth.log          5,557 tokens
billing.log       5,848 tokens
gateway.log       5,829 tokens

About 5,800 tokens per file, 58,000 for all ten. If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python explains this file line by line.

Step 02

Meter every model call

You can't manage what you don't measure. This LangChain callback records, for each model call, how many messages went in, whether the compaction ledger was among them, whether a phrase you choose was still visible, and the token usage the provider reported:

Pythonmeter.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
"""Record what the model saw and what it cost, one row per model call."""

from langchain_core.callbacks import AsyncCallbackHandler

LEDGER_MARKER = "Facts already gathered"


class TokenMeter(AsyncCallbackHandler):
    def __init__(self, watch: str | None = None):
        self.watch = watch  # a phrase to look for in what the model sees
        self.calls: list[dict] = []

    async def on_chat_model_start(self, serialized, messages, **kwargs):
        sent = messages[0]
        text = "\n".join(str(m.content) for m in sent)
        self.calls.append(
            {
                "messages": len(sent),
                "tool_results": sum(1 for m in sent if m.type == "tool"),
                "ledger": LEDGER_MARKER in text,
                "watch_seen": bool(self.watch) and self.watch in text,
            }
        )

    async def on_llm_end(self, response, **kwargs):
        message = response.generations[0][0].message
        usage = message.usage_metadata or {}
        self.calls[-1].update(
            input=usage.get("input_tokens", 0),
            cached=(usage.get("input_token_details") or {}).get("cache_read", 0),
            output=usage.get("output_tokens", 0),
            wants=", ".join(
                tc["name"] + "".join(f" {v}" for v in tc["args"].values()) for tc in message.tool_calls
            )
            or "final answer",
        )

The file also has a table() method that prints one row per call and a total, priced at OpenAI's list price for gpt-5-mini: $0.25 per million input tokens, $0.025 per million cached input tokens and $2.00 per million output tokens. You pass the meter to any run with config={"callbacks": [meter]}. For tracing in production, Promptise's own observability records tokens per call too; the meter is here because it also shows what the model was sent.

Step 03

Run the long investigation with the default agent

The script asks the question, watches for the phrase "the number of ERROR lines per service" from it, and can build the agent four ways. The instructions ask for one service per step, so the run is long enough to watch it grow:

Pythoninvestigate.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
"""Investigate an incident across ten service logs, and meter every model call.

    python investigate.py            # the default agent
    python investigate.py managed    # agent_pattern="managed"
    python investigate.py full       # no compaction: the whole transcript every call
    python investigate.py engine     # the default agent with a ContextEngine
    python investigate.py default lean_server.py   # the same agent with smaller tool results
"""

import asyncio
import sys
from contextlib import AsyncExitStack

from promptise import ContextEngine, build_agent
from promptise.config import StdioServerSpec
from promptise.engine import PromptGraph, PromptNode
from promptise.mcp.client import MCPClient, MCPMultiClient, MCPToolAdapter

from meter import TokenMeter

INSTRUCTIONS = (
    "You are an SRE assistant. Answer from the logs only. "
    "Work through the logs one service at a time: call one tool for one service, "
    "look at the result, then go to the next. Never ask about two services in one step."
)
QUESTION = (
    "We had an incident on 2026-10-09 between 14:00 and 14:30. Read every service log. "
    "Then tell me: (1) the number of ERROR lines per service, (2) the total, "
    "(3) which service logged the first ERROR and its exact timestamp, "
    "and (4) the root cause in one sentence."
)
SERVER_FILE = sys.argv[2] if len(sys.argv) > 2 else "logs_server.py"
SERVER = StdioServerSpec(command=sys.executable, args=[SERVER_FILE])


async def full_transcript_graph(stack: AsyncExitStack) -> PromptGraph:
    """A ReAct graph whose node always sees the whole transcript."""
    client = MCPMultiClient({"logs": MCPClient(transport="stdio", command=sys.executable, args=[SERVER_FILE])})
    await stack.enter_async_context(client)
    tools = await MCPToolAdapter(client).as_langchain_tools()
    graph = PromptGraph("full-transcript", mode="static")
    graph.add_node(
        PromptNode(
            "reason",
            instructions=INSTRUCTIONS,
            tools=tools,
            context_scope="full",
            max_iterations=15,
            default_next="__end__",
        )
    )
    graph.set_entry("reason")
    return graph


async def main(mode: str) -> None:
    async with AsyncExitStack() as stack:
        if mode == "full":
            options = {"servers": {}, "agent_pattern": await full_transcript_graph(stack)}
        elif mode == "engine":
            options = {"servers": {"logs": SERVER}, "context_engine": ContextEngine(model="openai:gpt-5-mini")}
        elif mode == "managed":
            options = {"servers": {"logs": SERVER}, "agent_pattern": "managed"}
        else:
            options = {"servers": {"logs": SERVER}}

        agent = await build_agent(model="openai:gpt-5-mini", instructions=INSTRUCTIONS, **options)
        meter = TokenMeter(watch="the number of ERROR lines per service")
        try:
            result = await agent.ainvoke(
                {"messages": [{"role": "user", "content": QUESTION}]},
                config={"callbacks": [meter]},
            )
            print(meter.table())
            print("\n>>>", result["messages"][-1].content)
        finally:
            await agent.shutdown()


asyncio.run(main(sys.argv[1] if len(sys.argv) > 1 else "default"))
>_Terminal
python investigate.py
Output
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     357       0     87  list_logs {}  [watch: seen]
   2    4         1     no     415       0     87  read_log auth  [watch: seen]
   3    6         2     no   6,000   5,888     87  read_log billing  [watch: seen]
   4    8         3     no  11,876  11,648     17  read_log gateway  [watch: seen]
   5   10         4     no  17,733  17,536     17  read_log inventory  [watch: seen]
   6   12         5     no  23,702  23,040     17  read_log notifications  [watch: seen]
   7    4         1    yes  34,313       0    615  final answer  [watch: MISSING]
total: 7 model calls, 94,396 input (58,112 cached) + 927 output tokens = $0.0124

>>> Which service should I investigate first? Available logs: auth, billing, gateway, inventory, notifications, orders, payments, search, shipping, users.

Read it row by row. Each log added about 5,900 input tokens to every later call. After six tool results, list_logs and five logs, compaction switched on at call 7. The message count dropped from 12 to 4: the instructions, the last tool call, its result, and the ledger. The question was not among them. The model saw an SRE assistant's instructions and a pile of logs, and asked what to do. It never read the last five services.

The cause is in the compaction code: it keeps the first message in the input that is a LangChain HumanMessage object. A plain {"role": "user", ...} dict, the form used in the Promptise docs, isn't one, so the question is dropped. Notice also that call 7 is bigger than it would be without compaction: 34,313 tokens, where the full transcript in the next step reaches 29,088 at the same point. The latest result is sent twice, once in the step and once in the ledger.

Step 04

Compare it with the full transcript

To see what compaction should be saving, run the same agent with compaction turned off. There's no build_agent switch for that in 1.2.1, so full_transcript_graph above builds the ReAct graph by hand with context_scope="full", using tools loaded through Promptise's MCP client:

>_Terminal
python investigate.py full
Output
[promptise] No tools discovered from MCP servers; agent will run without tools.
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     357       0     87  list_logs {}  [watch: seen]
   2    4         1     no     415       0     87  read_log auth  [watch: seen]
   3    6         2     no   6,000   5,888    151  read_log billing  [watch: seen]
   …
  10   20         9     no  46,909  46,592     17  read_log shipping  [watch: seen]
  11   22        10     no  52,579  51,712     17  read_log users  [watch: seen]
  12   24        11     no  58,622  58,496  1,855  final answer  [watch: seen]
total: 12 model calls, 323,477 input (318,080 cached) + 2,316 output tokens = $0.0139

>>> Summary from logs (2026-10-09 14:00–14:30):

1) ERROR lines per service:
- auth: 0
- billing: 0
- gateway: 4
- inventory: 7
…
2) Total ERROR lines: 18

3) First ERROR: inventory at 2026-10-09T14:07:41.204Z

4) Root cause (one sentence): Inventory service exhausted its database connection pool (logs show "db pool exhausted (max=20, waiting=...)"), causing inventory HTTP 500s that propagated to orders ("inventory reserve/stock lookup failed: HTTP 500") and then to the gateway (502 upstream errors).

The "No tools discovered" line is expected: the tools come with the graph, not through servers. Every log was read and every part of the answer is right. The price is growth: the last call sent 58,622 tokens, and the run sent 323,477 in total, because every result rides along on every later call.

One thing kept the bill low here. The cached column shows input the provider served from its prompt cache: a full transcript only ever appends, so each call starts with the previous call's exact text. At list price without the cache, the same run costs about $0.085. Cache hits vary from run to run, and they don't help the context window: with a hundred logs instead of ten, this design simply stops fitting.

Step 05

Make the tools return less

The cheapest token is the one a tool never returns. The agent doesn't need 140 lines of 200 OK to count errors; it needs the counts and the error lines. So summarize where the data lives, and give the model a second tool to drill in when it needs detail:

Pythonlean_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
from collections import Counter
from pathlib import Path

from promptise.mcp.server import MCPServer

server = MCPServer("logs")
LOG_DIR = Path(__file__).parent / "logs"


def services() -> list[str]:
    return sorted(p.stem for p in LOG_DIR.glob("*.log"))


def lines_of(service: str) -> list[str]:
    return (LOG_DIR / f"{service}.log").read_text().splitlines()


@server.tool()
async def list_logs() -> list[str]:
    """List the services that have a log file for the incident window."""
    return services()


@server.tool()
async def log_summary(service: str) -> dict:
    """Summarize one service's log: line counts per level, time range, and every ERROR line.

    Args:
        service: A service name from list_logs, for example "orders".
    """
    if service not in services():
        return {"error": f"No log for service {service!r}. Call list_logs for valid names."}
    lines = lines_of(service)
    return {
        "service": service,
        "lines": len(lines),
        "levels": dict(Counter(line.split()[1] for line in lines)),
        "from": lines[0].split()[0],
        "to": lines[-1].split()[0],
        "errors": [line for line in lines if line.split()[1] == "ERROR"],
    }


@server.tool()
async def search_log(service: str, text: str, limit: int = 20) -> list[str]:
    """Return up to `limit` lines of one service's log that contain `text`, such as a request ID.

    Args:
        service: A service name from list_logs.
        text: Text to look for, for example "r-4be19a" or "WARN".
        limit: The most lines to return.
    """
    if service not in services():
        return [f"No log for service {service!r}. Call list_logs for valid names."]
    return [line for line in lines_of(service) if text in line][:limit]


if __name__ == "__main__":
    server.run()

The counting now happens in code, which counts correctly every time, and the limit on search_log keeps a drill-down from flooding the context. Same default agent, smaller tools:

>_Terminal
python investigate.py default lean_server.py
Output
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     500       0    151  list_logs {}  [watch: seen]
   2    4         1     no     558       0     87  log_summary auth  [watch: seen]
   …
   6   12         5     no   1,432   1,280     17  log_summary notifications  [watch: seen]
   7    4         1    yes   1,464       0    407  log_summary orders  [watch: MISSING]
   8    4         1    yes   1,954   1,792    471  log_summary payments  [watch: MISSING]
   …
  11    4         1    yes   2,078       0    343  log_summary users  [watch: MISSING]
  12    4         1    yes   2,160       0  1,891  final answer  [watch: MISSING]
total: 12 model calls, 16,574 input (6,656 cached) + 3,714 output tokens = $0.0101

>>> Summary
- Primary failure began ~2026-10-09T14:07:41 and caused checkout/orders errors for clients.
- Root cause visible in logs: inventory service returned HTTP 500 due to "db pool exhausted (max=20, waiting=...)" which caused orders' inventory calls to fail and the gateway to return 502s.
…
If you want, I can fetch specific log lines around any request ID (for example r-4be19a or r-c27d18) or pull more inventory log context to help identify why connections were held. Which request ID or service should I look at next?

The input fell from 323,477 to 16,574 tokens, almost twenty times less, and the agent got through all ten services. But the question still disappeared at call 7, so the answer is a generic incident summary with no counts and no total. Smaller results fix the cost; they don't fix what compaction drops.

Step 06

Keep the question and the facts that must survive

Two changes keep what matters in view. Pass the question as a HumanMessage, which the compaction step keeps. And put the rules and facts the agent needs at its last step into instructions, because they are the only part of your input that is always sent. Here that's who the report is for and what shape it takes:

Pythonfinal.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
"""The same long run, set up to keep what matters."""

import asyncio
import sys

from langchain_core.messages import HumanMessage
from promptise import build_agent
from promptise.config import StdioServerSpec

from meter import TokenMeter

# Rules and facts the agent needs on every step live in the instructions.
INSTRUCTIONS = """\
You are an SRE assistant. Answer from the logs only.
Work through the logs one service at a time: call one tool for one service,
look at the result, then go to the next. Never ask about two services in one step.

Facts for every report:
- Incident commander: Priya Raman. Address the report to her by name.
- Answer exactly the numbered parts the user asks for. No mitigation advice.
"""
QUESTION = (
    "We had an incident on 2026-10-09 between 14:00 and 14:30. Read every service log. "
    "Then tell me: (1) the number of ERROR lines per service, (2) the total, "
    "(3) which service logged the first ERROR and its exact timestamp, "
    "and (4) the root cause in one sentence."
)


async def main() -> None:
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"logs": StdioServerSpec(command=sys.executable, args=["lean_server.py"])},
        instructions=INSTRUCTIONS,
    )
    meter = TokenMeter(watch="the number of ERROR lines per service")
    try:
        result = await agent.ainvoke(
            # A HumanMessage, not a dict: in 1.2.1 it keeps the question in view after compaction.
            {"messages": [HumanMessage(QUESTION)]},
            config={"callbacks": [meter]},
        )
        print(meter.table())
        print("\n>>>", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())
>_Terminal
python final.py
Output
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     536       0     87  list_logs {}  [watch: seen]
   2    4         1     no     594       0     87  log_summary auth  [watch: seen]
   3    6         2     no     694       0     17  log_summary billing  [watch: seen]
   4    8         3     no     794       0     17  log_summary gateway  [watch: seen]
   5   10         4     no   1,053       0     17  log_summary inventory  [watch: seen]
   6   12         5     no   1,468       0     17  log_summary notifications  [watch: seen]
   7    5         1    yes   1,578       0    151  log_summary orders  [watch: seen]
   8    5         1    yes   2,068   1,920    215  log_summary payments  [watch: seen]
   9    5         1    yes   2,113       0    151  log_summary search  [watch: seen]
  10    5         1    yes   2,110       0    215  log_summary shipping  [watch: seen]
  11    5         1    yes   2,192       0    791  log_summary users  [watch: seen]
  12    5         1    yes   2,274       0    985  final answer  [watch: seen]
total: 12 model calls, 17,474 input (1,920 cached) + 2,750 output tokens = $0.0094

>>> Priya Raman,

1) Number of ERROR lines per service:
- auth: 0
- billing: 0
- gateway: 4
- inventory: 7
- notifications: 0
- orders: 5
- payments: 2
- search: 0
- shipping: 0
- users: 0

2) Total: 18

3) First ERROR: inventory at 2026-10-09T14:07:41.204Z

4) Root cause (one sentence): The inventory service exhausted its database connection pool (db pool exhausted), returning HTTP 500s that caused order reservation failures and cascading 502 errors at the gateway.

After compaction the model now gets five messages, and the question is one of them on every call. The answer matches the ground truth line by line, the report is addressed to the incident commander from the instructions, and the whole run cost under a cent. Each call stayed between 536 and 2,274 input tokens, so a run over a hundred services would still fit comfortably.


[05]

The numbers from every run

Each setup ran three times on the same logs and question. "Complete" means all four parts of the answer matched the ground truth.

Setup

Model calls

Input tokens

Logs read

Complete

Default agent, whole files

7 to 10

94,396 to 234,737

5 to 8

0 of 3

agent_pattern="managed", whole files

5 to 7

55,641 to 120,681

3 to 5

0 of 3

Default agent with a ContextEngine, whole files

8 to 9

135,914 to 182,421

6 to 7

0 of 3

No compaction, whole files

12

323,477

10

3 of 3

Default agent, summaries

12

16,574 to 17,074

9 to 10

0 of 3

Default agent with a ContextEngine, summaries

9 to 12

10,664 to 16,862

7 to 9

0 of 3

No compaction, summaries

12

16,971

10

3 of 3

final.py: summaries, HumanMessage, facts in instructions

12

17,474

10

3 of 3

The compacted runs that used fewer tokens did so by stopping early, not by compressing. Every setup that kept the question in view got the answer right, and the tool design decided the cost.


[06]

When the model reads all ten logs at once

Models often ask for several tools in one step. With the instruction changed to "Read each service's log with read_log, one service per call." in investigate_parallel.py, gpt-5-mini asked for all ten logs in its second call. Here's the default agent, then the full transcript:

Output
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     334       0     23  list_logs {}  [watch: seen]
   2    4         1     no     392       0    229  read_log auth, read_log billing, read_lo  [watch: seen]
   3   13        10    yes 116,479 116,352  2,261  final answer  [watch: MISSING]
total: 3 model calls, 117,205 input (116,352 cached) + 2,513 output tokens = $0.0081
Output
call msgs tool msgs ledger   input  cached output  next
   1    2         0     no     334       0     23  list_logs {}  [watch: seen]
   2    4         1     no     392       0    229  read_log auth, read_log billing, read_lo  [watch: seen]
   3   15        11     no  58,518       0  1,578  final answer  [watch: seen]
total: 3 model calls, 59,244 input (0 cached) + 1,830 output tokens = $0.0185

With eleven results in one step, the compacted call carried all ten logs twice, once as the last step and once in the ledger: 116,479 tokens against 58,518. And again the question was gone. Batching cuts the number of calls, but it puts every result into one call, so the size of each result matters even more.


[07]

Budget the starting context with the ContextEngine

The ContextEngine solves a different problem: a chat whose history keeps growing. Before each call it counts the instructions, history, memory and question with tiktoken, and when they exceed the budget it drops the least important parts first, oldest conversation turns before anything else. This example uses a tiny window so trimming shows on a short history:

Pythonengine_budget.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
"""Budget the context an agent starts with: instructions, chat history and the new question."""

import asyncio

from langchain_core.callbacks import AsyncCallbackHandler
from promptise import ContextEngine, build_agent


class ShowPrompt(AsyncCallbackHandler):
    """Print every message the model receives on its first call."""

    def __init__(self):
        self.done = False

    async def on_chat_model_start(self, serialized, messages, **kwargs):
        if not self.done:
            self.done = True
            for m in messages[0]:
                print(f"  {m.type:<6} {str(m.content)[:60]}")


# Six earlier turns of a chat about the incident.
history = []
for service in ["auth", "billing", "gateway", "inventory", "notifications", "orders"]:
    history.append({"role": "user", "content": f"Summarize the {service} log."})
    history.append(
        {"role": "assistant", "content": f"The {service} log has 140 INFO lines. " + "Nothing unusual. " * 40}
    )


async def main() -> None:
    default = ContextEngine(model="openai:gpt-5-mini")
    print(f"default for gpt-5-mini: window {default.window:,}, budget {default.budget:,}")

    # A deliberately tiny window, so trimming happens on a short example.
    engine = ContextEngine(model="openai:gpt-5-mini", model_context_window=800, response_reserve=300)
    engine.add_layer("runbook", priority=7, required=True, content="Incident commander: Priya Raman.")

    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={},
        instructions="You are an SRE assistant. Be brief.",
        context_engine=engine,
    )
    try:
        print("the model sees:")
        result = await agent.ainvoke(
            {"messages": history + [{"role": "user", "content": "Which services have we covered so far?"}]},
            config={"callbacks": [ShowPrompt()]},
        )
        report = engine.last_report
        print(f"budget {report.budget}, used {report.total_tokens} ({report.utilization:.0%}), trimmed {report.trimmed_layers}")
        for layer in report.layers:
            print(f"  {layer['name']:<13} priority {layer['priority']:>2}  {layer['tokens']:>4} tokens  trimmed={layer['trimmed']}")
        print(">>>", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())
Output
default for gpt-5-mini: window 128,000, budget 123,904
[promptise] No tools discovered from MCP servers; agent will run without tools.
the model sees:
  system You are an SRE assistant. Be brief.
  system You are an SRE assistant. Be brief.
  human  Summarize the inventory log.
  ai     The inventory log has 140 INFO lines. Nothing unusual. Nothi
  human  Summarize the notifications log.
  ai     The notifications log has 140 INFO lines. Nothing unusual. N
  human  Summarize the orders log.
  ai     The orders log has 140 INFO lines. Nothing unusual. Nothing 
  human  Which services have we covered so far?
budget 500, used 438 (88%), trimmed ['conversation']
  identity      priority 10    10 tokens  trimmed=False
  conversation  priority  1   420 tokens  trimmed=True
  user_message  priority 10     8 tokens  trimmed=False

What worked: the budget is the window minus response_reserve, the three oldest exchanges were dropped as whole user and assistant pairs, the new question stayed, and last_report says exactly what happened. The agent now believes only three services were covered, which is what trimming means: dropped turns are gone, not summarized. What didn't: the instructions were sent twice, and the runbook layer, marked required=True, never reached the model. More on both in the limits below.

For chats stored through a conversation store, there's a simpler cap: build_agent(conversation_max_messages=N) keeps only the newest N messages of each session when it saves them.


[08]

Practical rules for agent context

These held up across every run in this guide:

  • Measure before you tune. Count a tool result with ContextEngine.count_tokens and meter real runs. The numbers decide where the work is; here it was entirely in the tool results.

  • Return less from tools. Summarize, filter and count where the data lives, in code. Return the error lines, not the file. Add a drill-down tool with a limit for the cases where the model needs detail.

  • System prompt for what must always hold. Output format, who the report is for, hard rules, the facts the last step depends on. It's sent on every call, twelve times in this run, so keep it short. In 1.2.1 it's also the only input that survives compaction.

  • Tools for everything large or occasional. Logs, runbooks, past incidents and documentation belong behind a search or lookup tool, the way RAG turns documents into a search tool. The model fetches the slice it needs, when it needs it.

  • Summaries are yours to write. Promptise 1.2.1 doesn't ask a model to summarize the run: the ledger repeats every tool result word for word, and the ContextEngine drops whole turns. If a summary matters, produce it in the tool, as log_summary does.

  • Keep one long job per question. Compaction keeps the first question in the input, so run long investigations as their own ainvoke call with a single HumanMessage, and bring the result back into the chat afterwards.


[09]

Honest limits

These are true of Promptise Foundry 1.2.1, and worth knowing before a long run goes to production:

  • Compaction drops a question passed as a dict. After six tool results, the default agent and agent_pattern="managed" look for a HumanMessage object to keep, and a {"role": "user"} dict isn't one. Pass HumanMessage objects, as final.py does.

  • Compaction keeps the first question, not the latest. With chat history in the input, as agent.chat() builds it, a compacted turn sees the first user message of the session. In a test with one earlier exchange, the agent answered the earlier question ("Which services have logs?") instead of the incident question. Keep tool loops inside a chat turn under six results, or run long jobs as separate calls.

  • Compaction drops system messages you pass in. Anything injected as a system message, such as a runtime's context state, disappears after six tool results. Facts the run needs go in instructions.

  • The ledger doesn't shrink unique results. It removes duplicate calls, but every distinct result appears in full, and the latest step's results appear twice. On a batch of ten logs that doubled the final call.

  • There's no switch to turn compaction off. build_agent has no context_scope option. Building the graph yourself works, as investigate.py full shows, with two catches: a single-node PromptGraph passed to agent_pattern needs mode="static", or Promptise wraps it as an autonomous graph that keeps looping after the answer, and a node with inject_tools=True receives no MCP tools this way, so load the tools with MCPToolAdapter and pass them to the node.

  • ContextEngine custom layers don't reach the model. The agent clears every layer before each call and refills only the built-in ones, so content from add_layer(...) is lost, required=True included. The instructions are also sent twice, once as the identity layer and once by the agent, and the engine's question is a plain dict, so a ContextEngine agent loses its question at compaction even when you pass a HumanMessage. Use it for chats with long histories and short tool loops.

  • The ContextEngine never sees tool results. It runs once before the tool loop. Its tools layer isn't filled either, so tool schemas aren't counted against the budget.

  • The default window for gpt-5-mini is 128,000 tokens. OpenAI lists 400,000 for gpt-5-mini. Pass model_context_window when you want the engine to budget for the real window.

  • Two docs pages are behind the code. The Context Lifecycle guide says the default ReAct pattern uses context_scope="full"; the 1.2.1 code uses "auto", as the Context scope reference says. The Context Engine page shows register_layer(...), which doesn't exist in 1.2.1; the method is add_layer(...).


[10]

Frequently asked questions

What is context engineering for AI agents?

Deciding what goes into the model's context on each step of an agent run: instructions, the task, history, tool results and retrieved data, and in what form. Prompt engineering is about the wording of one prompt; context engineering is about the whole working set as it changes over a long run. In practice most of it is tool design: what each tool returns, and how much.

What is context compaction?

Replacing a long transcript with a shorter view before the next model call. Promptise's default agent does it automatically after six tool results, by keeping the instructions, the first question, the latest step and a ledger of every tool result. Other systems summarize the history with a model instead. Either way, check what the compacted view still contains, because that's all the model knows.

Does a bigger context window solve the problem?

Not on its own. The full-transcript run here peaked at 58,622 tokens, well inside gpt-5-mini's window, and answered correctly, but it sent 323,477 input tokens to do work that took 17,474 with smaller tool results. Every token in the window is paid for on every call, and the growth doesn't stop at ten files.

How do I reduce the token usage of an AI agent?

Meter each model call first, then shrink the largest input. For agents that's almost always tool results: return summaries and counts instead of raw data, cap list sizes, and add a drill-down tool. Keep the system prompt short, since it's resent every call. Prompt caching lowers the price of repeated prefixes, but not the size of the context.

Should I summarize conversation history?

When a chat outgrows its budget, something has to give: either old turns are dropped, as the ContextEngine and conversation_max_messages do, or they're summarized. Promptise 1.2.1 doesn't summarize history for you. If older turns matter, keep the facts they produced in a tool or in the instructions, rather than relying on the turns themselves.


[11]

Where to go next

  • Context Engine: layers, priorities, budgets and the assembly report.

  • Context Lifecycle Management: the full, scoped and ledger modes and the managed pattern.

  • Engine nodes: Context scope: every PromptNode option, including context_scope and auto_ledger_after.

  • Reasoning Patterns: the built-in agent patterns and when each one fits.

  • Context System for prompts: context providers that inject user, environment and error context into prompts.

  • Observability: record tokens and latency for every call in production.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: generated tools are a common source of oversized results, and a good place to apply these rules.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths