Most agents are written as a request and a reply: you call them, they answer, they're gone. The agents that do real work in production look different. They wake up on a schedule or when something happens, do one bounded job, remember what matters for next time, and stop themselves before they do too much. This guide builds one of those in Python with the Promptise Agent Runtime: a support-ticket triage agent that runs every few minutes, keeps state between runs, writes a journal you can read back, and pauses itself when it goes over its budget. Every snippet ran against Promptise Foundry 1.2.1, and the output you see is what those runs printed, including the parts that didn't go to plan.
How do you deploy AI agents in production?
Run the agent as a long-lived process, not a function call. Give it a trigger that wakes it (a cron schedule, a webhook, a file landing in a folder), tools that reach your systems, limits on what one run may do, and a record of what it did. With Promptise Foundry, that process is an AgentProcess:
import asyncio
import sys
from promptise.config import StdioServerSpec
from promptise.runtime import AgentProcess, ProcessConfig, TriggerConfig
async def main() -> None:
process = AgentProcess(
name="triage",
config=ProcessConfig(
model="openai:gpt-5-mini",
instructions="Triage new support tickets, then post one short summary.",
servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
triggers=[TriggerConfig(type="cron", cron_expression="*/15 * * * *")],
),
)
await process.start() # builds the agent, connects its tools, starts the trigger
print(process.status()["state"], "with", process.status()["trigger_count"], "trigger")
try:
await asyncio.Event().wait() # from here on, the agent wakes every 15 minutes
finally:
await process.stop()
print(process.status()["state"])
if __name__ == "__main__":
try:
asyncio.run(main())
except KeyboardInterrupt:
passrunning with 1 trigger
stoppedThat's a working AI agent cron job. start() builds the agent, connects to its MCP tools and starts the trigger. From then on the cron trigger wakes it every 15 minutes, with no ainvoke() call from you, until you stop it. Everything else in this guide is what you add before you trust it to run unattended.
[02]
How a long-running agent works
An AgentProcess wraps an ordinary Promptise agent. Triggers put events on a queue. A worker takes each event, adds what the agent should know (its context state, recent conversation, remaining budget) and runs the agent once. After the run, the runtime checks the budget and decides whether the process keeps going.
Rendering diagram…
Each process moves through a small state machine, and every transition is validated:
State | How the process gets there |
|---|---|
running | start(), or resume() after a pause |
suspended | suspend(), a budget with on_exceeded="pause", or a health anomaly with on_anomaly="pause" |
stopped | stop(), or a budget with on_exceeded="stop" |
failed | A startup error, or max_consecutive_failures failed runs in a row (default 3) |
A suspended process keeps its triggers. Events that fire while it's paused wait in the queue until someone calls resume().
[03]
What you need
Python 3.10 or newer.
Promptise Foundry: pip install promptise. Cron, webhook and file-watch triggers work out of the box; their dependencies ship with the base install.
An API key for a model provider. This guide uses OpenAI with openai:gpt-5-mini; set OPENAI_API_KEY in your environment or a .env file.
[04]
Build a ticket triage agent, step by step
The agent's job: every time it wakes, find new support tickets, give each a priority, and post one summary to the team channel, naming the on-call engineer for anything urgent.
Give the agent its tools
The agent reaches your systems through MCP tools. Here a small MCP server stands in for a ticketing system and a team chat: a JSON file of tickets, and a log file as the channel. If you haven't built one before, How to Connect MCP Servers to Your AI Agent in Python covers the basics.
import json
from datetime import datetime
from pathlib import Path
from typing import Literal
from promptise.mcp.server import MCPServer, ToolError
server = MCPServer("tickets")
# A stand-in for your ticketing system: a JSON file the agent reads and updates.
TICKETS = Path(__file__).with_name("tickets.json")
CHANNEL = Path(__file__).with_name("support-channel.log")
def load() -> dict:
return json.loads(TICKETS.read_text()) if TICKETS.exists() else {}
@server.tool(read_only_hint=True)
async def list_untriaged() -> list[dict]:
"""List new support tickets that don't have a priority yet."""
return [{"ticket_id": tid, **t} for tid, t in load().items() if t.get("priority") is None]
@server.tool()
async def set_priority(ticket_id: str, priority: Literal["P1", "P2", "P3"], reason: str) -> dict:
"""Give a ticket a priority. P1 = outage or data loss, P2 = a customer is blocked, P3 = everything else.
Args:
reason: One short sentence explaining the priority.
"""
tickets = load()
if ticket_id not in tickets:
raise ToolError(f"There is no ticket {ticket_id}.")
tickets[ticket_id].update(priority=priority, reason=reason)
TICKETS.write_text(json.dumps(tickets, indent=2))
return {"ticket_id": ticket_id, "priority": priority}
@server.tool()
async def post_summary(text: str) -> str:
"""Post a short triage summary to the support team's channel."""
stamp = datetime.now().strftime("%H:%M:%S")
with CHANNEL.open("a") as f:
f.write(f"[{stamp}] {text}\n")
return "posted"
if __name__ == "__main__":
server.run()Notice where the durable state lives: in the ticket system. A ticket with a priority is done, so the next run won't touch it again. That's the most reliable memory an agent can have, and it's yours, not the runtime's.
Describe the process
ProcessConfig holds everything about the process: model, instructions, tools, triggers, starting state and limits. The file also writes a journal, which Step 6 explains.
import asyncio
import sys
from promptise import CallbackSink, EventNotifier
from promptise.config import StdioServerSpec
from promptise.runtime import (
AgentProcess,
BudgetConfig,
ContextConfig,
FileJournal,
JournalEntry,
ProcessConfig,
ToolCostAnnotation,
TriggerConfig,
)
INSTRUCTIONS = """\
You triage the support queue. Each time you wake up:
1. Call list_untriaged.
2. Give every ticket a priority with set_priority.
3. If you set any priorities, post ONE short summary with post_summary.
For every P1, name the on-call engineer: use the exact on_call value
from [Context State]. Never invent a name.
If there are no new tickets, don't post anything and reply "Nothing to triage."
"""
CONFIG = ProcessConfig(
model="openai:gpt-5-mini",
instructions=INSTRUCTIONS,
servers={
"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"]),
},
# Every minute, for the demo. In production: "*/15 * * * *" or "0 8 * * 1-5".
triggers=[TriggerConfig(type="cron", cron_expression="* * * * *")],
# State the agent sees at the start of every run.
context={"initial_state": {"on_call": "Sam"}},
# The autonomy budget: what one run, and one day, may do.
budget=BudgetConfig(
enabled=True,
max_tool_calls_per_run=6,
max_runs_per_day=200,
tool_costs={"post_summary": ToolCostAnnotation(cost_weight=1.0, irreversible=True)},
max_irreversible_per_run=1,
on_exceeded="pause",
),
)
journal = FileJournal(".promptise/journal")
def build_process(name: str = "triage", initial_state: dict | None = None) -> AgentProcess:
config = CONFIG
if initial_state is not None:
config = CONFIG.model_copy(update={"context": ContextConfig(initial_state=initial_state)})
async def record(event) -> None:
"""Write every runtime event to the journal, plus a checkpoint after each run."""
await journal.append(JournalEntry(process_id=name, entry_type=event.event_type, data=event.data))
if event.event_type == "invocation.complete":
await journal.checkpoint(
name,
{
"context_state": process.context.state_snapshot(),
"lifecycle_state": process.state.value,
"budget_remaining": process.status().get("budget"),
},
)
process = AgentProcess(
name=name,
process_id=name,
config=config,
event_notifier=EventNotifier(sinks=[CallbackSink(record)]),
)
return process
async def main() -> None:
process = build_process()
await process.start()
print(f"{process.name} is {process.state.value}. Press Ctrl+C to stop.")
try:
while True:
await asyncio.sleep(3600)
finally:
await process.stop()
print(f"{process.name} is {process.state.value}.")
if __name__ == "__main__":
try:
asyncio.run(main())
except KeyboardInterrupt:
passThe parts that matter:
The trigger. A standard five-field cron expression. This guide uses every minute so you can watch several cycles; a real triage agent might run every 15 minutes, or at 8:00 on weekdays.
The context. initial_state is the process's key-value state. The runtime shows it to the agent at the start of every run, and your code can change it while the process runs.
The budget. At most six tool calls per run, 200 runs a day, and one irreversible action per run. post_summary is marked irreversible because a message, once posted, can't be unposted. Cost weights are abstract units you choose, not money.
python triage_agent.py runs the process until you press Ctrl+C:
triage is running. Press Ctrl+C to stop.
triage is stopped.Run it for a few cycles
To see the whole lifecycle in a few minutes, this script starts the process, adds tickets between runs and prints what happens. A small helper, seed_tickets.py, adds the tickets: two at first, one between runs, then a burst of five.
"""Drive the triage process through three cron cycles and print what happens."""
import asyncio
import subprocess
import sys
import time
from pathlib import Path
from triage_agent import build_process
CHANNEL = Path("support-channel.log")
def seed(batch: str) -> None:
out = subprocess.run([sys.executable, "seed_tickets.py", batch], capture_output=True, text=True)
print(f"{time.strftime('%H:%M:%S')} queue: {out.stdout.strip()}")
async def wait_for_run(process, n: int) -> None:
while process.status()["invocation_count"] < n:
await asyncio.sleep(1)
await asyncio.sleep(1) # let post-run bookkeeping finish
s = process.status()
print(
f"{time.strftime('%H:%M:%S')} run {n} done: state={s['state']} "
f"conversation_messages={s['conversation_messages']} "
f"tool_calls_left_this_run={s['budget']['tool_calls_run']}"
)
async def main() -> None:
CHANNEL.unlink(missing_ok=True)
seed("1")
process = build_process()
await process.start()
print(f"{time.strftime('%H:%M:%S')} {process.name} is {process.state.value}")
await wait_for_run(process, 1)
# Shift change: tell the running agent who is on call now.
process.context.put("on_call", "Dana", source="operator")
seed("2")
await wait_for_run(process, 2)
# A burst of tickets: more work than one run's budget allows.
seed("burst")
await wait_for_run(process, 3)
# Wait for the next cron tick while paused.
await asyncio.sleep(65)
s = process.status()
print(f"{time.strftime('%H:%M:%S')} after next tick: state={s['state']} queue_size={s['queue_size']}")
print("\nsupport-channel.log:")
print(CHANNEL.read_text())
print("on_call history:", [(e.value, e.source) for e in process.context.state_history("on_call")])
await process.stop()
print(f"{time.strftime('%H:%M:%S')} {process.name} is {process.state.value}")
asyncio.run(main())19:06:33 queue: 2 new ticket(s) in the queue
19:06:56 triage is running
19:07:12 run 1 done: state=running conversation_messages=2 tool_calls_left_this_run=2
19:07:12 queue: 1 new ticket(s) in the queue
19:08:09 run 2 done: state=running conversation_messages=4 tool_calls_left_this_run=3
19:08:09 queue: 5 new ticket(s) in the queue
Budget violation: max_tool_calls_per_run (limit=6, current=7) — action=pause
19:09:16 run 3 done: state=suspended conversation_messages=6 tool_calls_left_this_run=-1
19:10:21 after next tick: state=suspended queue_size=1
support-channel.log:
[19:07:09] P1: T-201 — Checkout 500s since 09:10 (on-call: Sam)
P2: T-202 — Customer can't download invoice PDF, blocking billing.
[19:08:06] P1: T-203 — Nobody in the EU region can log in. On-call: Dana.
[19:09:10] Triage: set priorities — T-300: P2 (password reset email not received); T-301: P2 (mobile app crashes opening settings); T-302: P2 (double charge); T-303: P3 (dark mode logo invisible); T-304: P2 (enterprise customer receiving 429). No P1s.
on_call history: [('Sam', 'system'), ('Dana', 'operator')]
19:10:21 triage is stoppedThree runs, each at the top of a minute, with no call from the script. Starting the process took about 20 seconds here; start() builds the agent and connects its tools once, and every run reuses them. The next three steps read this output from top to bottom.
Change what the agent knows while it runs
Between runs 1 and 2, the on-call engineer changed. The script told the running process with one line:
process.context.put("on_call", "Dana", source="operator")Run 1 named Sam for its P1. Run 2 named Dana. Nothing restarted. The context keeps an audit trail too: state_history("on_call") shows who set each value, the initial system value and the operator change.
The process carries two kinds of memory between runs:
Context state, the key-value store you just used. The runtime adds it to the start of every run as a [Context State] message.
Recent conversation, a rolling buffer of each run's trigger message and the agent's final reply. That's why conversation_messages grows by two per run. The default keeps the last 100 messages; change it with conversation_max_messages in ContextConfig.
Here's what the model received on each call of that test, captured with a LangChain callback:
MODEL SEES: system | You triage the support queue. Each time you wake up: 1. Call list_untriaged. 2. Give every ticket a priorit…
MODEL SEES: system | [Context State] {'on_call': 'Dana'}
MODEL SEES: human | [Trigger: cron] Payload: {}
---
…
---
MODEL SEES: system | You triage the support queue. Each time you wake up: 1. Call list_untriaged. 2. Give every ticket a priorit…
MODEL SEES: ai |
MODEL SEES: tool | {"ticket_id": "T-300", "priority": "P2"}
…
MODEL SEES: system | Facts already gathered (do NOT call a tool to fetch any of these again): - list_untriaged({'payload': {}}) …Let the budget stop a runaway run
Run 3 got five tickets at once. Listing them, prioritizing all five and posting the summary takes seven tool calls; the budget allows six. Here's what happened, from the output above:
The run finished. All five tickets got a priority and the summary was posted.
After the run, the runtime saw the seventh call over the limit, logged
Budget violation: max_tool_calls_per_run (limit=6, current=7) — action=pause, and suspended the process.The next cron tick still fired. Its event waited in the queue (queue_size=1) instead of running.
That's the most important thing to know about the autonomy budget in 1.2.1: per-run limits are checked when the run ends, not in the middle of it. The run that crosses the line completes; the budget stops the next one. It's a circuit breaker for runaway schedules, not a hard cap on a single run. The daily limit, max_runs_per_day, is the exception: it's checked before a run starts, so the run over the limit never happens.
So set per-run limits a little above what a normal run needs, and use the daily limits to cap the total. Three more things worth knowing:
on_exceeded can be "pause", "stop" or "escalate". "escalate" posts the violation to a webhook or an event bus, then pauses.
Nothing resumes a paused process automatically, not even the daily reset. A person, or your own code, calls process.resume(), and the queued events run then.
With inject_remaining=True (the default), the agent sees its remaining budget at the start of each run as a [Budget Remaining] message.
Keep a journal you can read back
When something goes wrong at 3 a.m., you want a record of what the agent did. Promptise has a journal for exactly this: FileJournal writes append-only JSONL, one file per process, and promptise runtime logs prints it.
The process emits events as it works: process.started, invocation.start, invocation.complete, budget.warning, budget.exceeded and more. Pass an EventNotifier with a CallbackSink, and you can journal every one of them. This is the record function from Step 2:
async def record(event) -> None:
"""Write every runtime event to the journal, plus a checkpoint after each run."""
await journal.append(JournalEntry(process_id=name, entry_type=event.event_type, data=event.data))
if event.event_type == "invocation.complete":
await journal.checkpoint(
name,
{
"context_state": process.context.state_snapshot(),
"lifecycle_state": process.state.value,
"budget_remaining": process.status().get("budget"),
},
)After each run it also stores a checkpoint: a snapshot of the context state, the lifecycle state and the remaining budget. Read the journal back from the command line:
promptise runtime logs triage --lines 40 Journal: triage
┏━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Time ┃ Type ┃ Data ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ 2026-10-10T17:06:56.963 │ process.started │ {"process_name": "triage", │
│ │ │ "process_id": "triage"} │
│ 2026-10-10T17:07:00.118 │ invocation.start │ {"model": │
│ │ │ "openai:gpt-5-mini"} │
│ 2026-10-10T17:07:10.289 │ invocation.complete │ {"duration_ms": 10223.2} │
│ 2026-10-10T17:07:10.290 │ checkpoint │ {"state": {"context_state": │
│ │ │ {"on_call": "Sam"}, │
│ │ │ "lifecycle_state": │
│ │ │ "running",... │
…
│ 2026-10-10T17:09:14.947 │ invocation.complete │ {"duration_ms": 14936.9} │
│ 2026-10-10T17:09:14.948 │ checkpoint │ {"state": {"context_state": │
│ │ │ {"on_call": "Dana"}, │
│ │ │ "lifecycle_state": │
│ │ │ "suspende... │
│ 2026-10-10T17:09:14.956 │ budget.exceeded │ {"process_name": "triage", │
│ │ │ "limit_type": │
│ │ │ "max_tool_calls_per_run", │
│ │ │ "current":... │
└─────────────────────────┴──────────────────────┴─────────────────────────────┘The full entries are in .promptise/journal/triage.jsonl, for example this line from run 3:
{"entry_id": "…", "process_id": "triage", "timestamp": "2026-10-10T17:09:14.956…", "entry_type": "budget.exceeded", "data": {"process_name": "triage", "limit_type": "max_tool_calls_per_run", "current": 7, "limit": 6}}promptise runtime logs looks up the journal by the name you give it, which is why build_process uses the process name as its process_id too.
Restart without losing state
The process keeps its context in memory, so a restart starts from the initial state again. The checkpoint fixes that. ReplayEngine reads the journal and returns the last known state, and you start the new process with it:
import asyncio
from promptise.runtime import ReplayEngine
from triage_agent import build_process, journal
async def main() -> None:
recovered = await ReplayEngine(journal).recover("triage")
print("recovered:", recovered)
process = build_process(initial_state=recovered["context_state"])
await process.start()
print("state after restart:", process.context.state_snapshot())
await process.stop()
asyncio.run(main())recovered: {'context_state': {'on_call': 'Dana'}, 'lifecycle_state': 'suspended', 'last_entry_type': 'budget.warning', 'entries_replayed': 10}
state after restart: {'on_call': 'Dana'}The new process knows Dana is on call, although the config still says Sam. It also learned the old process was suspended when it stopped: here, paused by its budget. Starting it again is a decision for a person, so check lifecycle_state before you do.
[05]
Wake on a webhook instead of a clock
A schedule is right for sweeps. When something else already knows a ticket arrived, let it wake the agent. A webhook trigger starts a small HTTP server inside the process; this keeps a slow cron sweep as a safety net:
import asyncio
from promptise.runtime import AgentProcess, TriggerConfig
from triage_agent import CONFIG
config = CONFIG.model_copy(
update={
"triggers": [
TriggerConfig(type="cron", cron_expression="*/15 * * * *"),
TriggerConfig(type="webhook", webhook_path="/tickets", webhook_port=8180),
]
}
)
async def main() -> None:
process = AgentProcess(name="triage", config=config)
await process.start()
print("listening on http://127.0.0.1:8180/tickets")
while process.status()["invocation_count"] < 1:
await asyncio.sleep(1)
print("runs:", process.status()["invocation_count"])
await process.stop()
asyncio.run(main())Post an event the way your help desk would:
curl -s -X POST http://127.0.0.1:8180/tickets -H 'Content-Type: application/json' -d '{"event": "ticket.created", "ticket_id": "T-401"}'{"status": "accepted", "event_id": "494fb90b-8bd7-4f7e-a816-bef0093f1e0e"}The webhook answers 202 at once and the run happens in the background. The JSON body becomes the trigger payload the agent sees. The process printed:
WebhookTrigger on port 8180 has no HMAC secret — any HTTP client can trigger this webhook. Set hmac_secret for production use.
listening on http://127.0.0.1:8180/tickets
runs: 1And the channel got:
[19:13:16] P1: T-401 — Data export deletes rows from source table (possible data loss). On-call: Sam.Take that warning seriously. A webhook created from TriggerConfig listens on 127.0.0.1 only and has no signature check, and TriggerConfig has no field to change either. The WebhookTrigger class itself takes host and hmac_secret, so register a trigger type that builds it with a secret:
register_trigger_type(
"signed_webhook",
lambda config, **_: WebhookTrigger(path=config.webhook_path, port=config.webhook_port,
hmac_secret=config.custom_config["secret"]),
)Use it as TriggerConfig(type="signed_webhook", ..., custom_config={"secret": ...}). In a test, an unsigned POST got 401 and a POST with a valid X-Webhook-Signature: sha256=<hex> header got 202.
The other built-in triggers work the same way. file_watch wakes the agent when files matching watch_patterns appear or change in watch_path. event and message listen to an event bus or message broker, which you pass to an AgentRuntime to share between its processes. Triggers lists them all, and how to register your own.
[06]
Declare it in a manifest and run it from the CLI
You don't have to write Python to run a process. A .agent manifest describes the same thing in YAML, and the promptise runtime command runs it:
version: "1.0"
name: triage
model: openai:gpt-5-mini
instructions: |
You triage the support queue. Each time you wake up:
1. Call list_untriaged.
2. Give every ticket a priority with set_priority.
3. If you set any priorities, post ONE short summary with post_summary.
For every P1, name the on-call engineer: use the exact on_call value
from [Context State]. Never invent a name.
If there are no new tickets, don't post anything and reply "Nothing to triage."
servers:
tickets:
type: stdio
command: python
args: ["tickets_server.py"]
triggers:
- type: cron
cron_expression: "* * * * *" # every minute for the demo
world:
on_call: Sam
budget:
enabled: true
max_tool_calls_per_run: 6
max_runs_per_day: 200
max_irreversible_per_run: 1
on_exceeded: pause
tool_costs:
post_summary:
cost_weight: 1.0
irreversible: trueworld is the manifest's name for the initial context state. Check the file before you run it:
promptise runtime validate triage.agentValidating triage.agent...
✓ Schema validation passed
✓ No warnings
Manifest Summary
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Field ┃ Value ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ Name │ triage │
│ Model │ openai:gpt-5-mini │
│ Version │ 1.0 │
│ Instructions │ You triage the support queue. Each time you wake up: │
│ │ 1. Call... │
│ Servers │ 1 │
│ Triggers │ 1 │
└──────────────┴──────────────────────────────────────────────────────┘
✓ Validation complete!Then start it. It runs in the foreground until you press Ctrl+C:
promptise runtime start triage.agentLoaded process: triage
Agent Processes
┏━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┓
┃ Name ┃ State ┃ PID ┃ Invocations ┃ Queue ┃ Uptime ┃
┡━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━┩
│ triage │ RUNNING │ 51495264 │ 0 │ 0 │ 0h 0m 0s │
└────────┴─────────┴──────────┴─────────────┴───────┴──────────┘
Press Ctrl+C to stop all processes.
Shutting down...
All processes stopped.In that run the agent woke at the next minute, triaged two tickets and posted its summary before Ctrl+C. Pass a directory instead of a file to start every .agent file in it, and promptise runtime init --template cron writes a starter manifest. The Agent Manifests and Runtime CLI docs have every field and option.
[07]
Other guardrails: health, mission and secrets
The budget limits how much an agent does. Three more opt-in subsystems watch what it does. All are off by default.
Behavioral health (HealthConfig) spots bad patterns without calling a model: the same tool with the same arguments again and again, repeating loops, empty responses and high error rates. on_anomaly can log, pause or escalate.
Mission (MissionConfig) gives a process an objective and success_criteria. Every eval_every runs, a model judges progress; the process can stop itself when the mission is achieved, or escalate when confidence drops below confidence_threshold.
Secret scoping (SecretScopeConfig) gives each process its own secrets, resolved from ${ENV_VAR} references when it starts, encrypted in memory, with optional expiry, and revoked on stop(). A missing variable stops the start with
KeyError: "Environment variable 'NOT_SET_ANYWHERE' not found (required by secret config)". The agent itself can only read them through the get_secret tool in open execution mode.
Tune health checks before you let them pause anything. With the default thresholds, an agent whose queue was empty three runs in a row was paused as unhealthy, because three short tool results in a row ([]) count as empty responses:
run 1: state=running anomalies=0
run 2: state=running anomalies=0
Behavioral anomaly detected: empty_response — Agent producing empty responses: 3 consecutive responses under 10 characters
run 3: state=suspended anomalies=1Start with on_anomaly="log", read what it flags for a few days, then raise empty_threshold and stuck_threshold to fit your agent before switching to "pause". See Behavioral Health, Mission Model and Secret Scoping.
[08]
Before you go to production
Know where every piece of state lives. In 1.2.1:
State | Where it lives | Survives a restart? |
|---|---|---|
Context state | Process memory | Only if you checkpoint and restore it, as in Step 7 |
Recent conversation | Process memory; stop() clears it | No |
Budget counters, daily ones too | Process memory | No: they start from zero |
Journal and checkpoints | Files you write | Yes |
Tickets, orders, records | Your systems, behind MCP tools | Yes |
With that in mind:
Restarts are your supervisor's job. ProcessConfig has restart_policy and max_restarts, but 1.2.1 doesn't act on them. In a test with restart_policy="on_failure", a process that failed three runs in a row went to failed and was still failed 15 seconds later. Run each process under systemd, Docker or Kubernetes with a restart policy, restore context from your checkpoint on start, and alert on failed.
Run one copy of each process. Each process has its own cron trigger and nothing coordinates replicas, so two copies of the same process both fire at every tick and do the work twice. Scale by giving different processes different jobs, and make tools that change data safe to repeat.
Distributed means remote control, not failover. RuntimeTransport puts an HTTP API on a node's runtime (health, status, start, stop, inject an event), with an optional bearer token. RuntimeCoordinator registers nodes, checks their health and aggregates status. There's no shared state, no leader election and no automatic move of processes off a failed node. Also, the coordinator can't send a token: against a node with auth_token set, its health check still worked, but get_node_status failed with
Status request failed: 401.Keep runs small. Short runs stay inside the budget, keep the context state in view (see Step 4), and are cheaper to retry.
Lock down triggers. Webhooks need a signature check and a reverse proxy in front if anything outside the machine calls them.
Watch the journal. Alert on budget.exceeded, health.anomaly and invocation.error events; the CallbackSink from Step 6 is a good place to send them on.
[09]
Frequently asked questions
How do I run an AI agent as a cron job?
Give an AgentProcess a trigger like TriggerConfig(type="cron", cron_expression="*/15 * * * *") and start it. The process stays up and wakes the agent on schedule, so the agent and its tool connections are built once instead of on every run. A plain system cron job that runs a script also works; it just rebuilds everything each time and keeps no state in between.
How does a long-running AI agent remember things between runs?
Three ways, in order of reliability. Your own systems, through tools: a triaged ticket stays triaged. The process context, a key-value store the agent sees at the start of every run, which you can checkpoint to a journal and restore. And a rolling buffer of recent trigger messages and replies. For anything that must survive a restart, use the first two.
How do I stop an autonomous AI agent from running up costs?
Enable BudgetConfig with per-run and daily limits on tool calls, LLM turns, cost weights and irreversible actions, and choose pause, stop or escalate for when a limit is crossed. Per-run limits take effect after the run that crosses them; max_runs_per_day is checked before a run starts, so it's your hard cap. Track real token spend with your model provider as well.
Can I run AI agents on several machines?
You can run a runtime on each machine and manage them over HTTP with RuntimeTransport and RuntimeCoordinator. In 1.2.1 that gives you remote start, stop, status and health checks, not shared state or automatic failover, so assign each process to one machine.
Do I need Kubernetes to deploy an AI agent?
No. One promptise runtime start in the foreground under any supervisor that restarts it on failure is enough to begin with: systemd, Docker with a restart policy, or a Kubernetes Deployment with one replica.
[10]
Where to go next
Agent Runtime: the overview of processes, triggers, state and governance.
Agent Processes: every ProcessConfig field and lifecycle method.
Cron triggers and Triggers: schedules, webhooks, file watches and custom trigger types.
Journal and Context & State: what to record and what the agent sees.
Runtime Manager: several processes in one AgentRuntime.
Build a Production MCP Server in Python: make the tools this agent calls safe for unattended use.
OpenAPI to MCP: Turn Any REST API into an MCP Server: give the agent tools for an API you already have.