PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

Deploy AI Agents in Production: Long-Running Python Agents

Run an AI agent as a long-lived Python process: cron and webhook triggers, state between runs, a journal, budgets that pause it, and honest limits.

Level
Intermediate
Reading time
15 min
Published
Oct 10, 2026
By
Promptise Team
  • AI Agents
  • Agent Runtime
  • Production
  • Cron
  • Python
  • Promptise Foundry

Most agents are written as a request and a reply: you call them, they answer, they're gone. The agents that do real work in production look different. They wake up on a schedule or when something happens, do one bounded job, remember what matters for next time, and stop themselves before they do too much. This guide builds one of those in Python with the Promptise Agent Runtime: a support-ticket triage agent that runs every few minutes, keeps state between runs, writes a journal you can read back, and pauses itself when it goes over its budget. Every snippet ran against Promptise Foundry 1.2.1, and the output you see is what those runs printed, including the parts that didn't go to plan.

[01]

How do you deploy AI agents in production?

Run the agent as a long-lived process, not a function call. Give it a trigger that wakes it (a cron schedule, a webhook, a file landing in a folder), tools that reach your systems, limits on what one run may do, and a record of what it did. With Promptise Foundry, that process is an AgentProcess:

Pythonquickstart.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
import asyncio
import sys

from promptise.config import StdioServerSpec
from promptise.runtime import AgentProcess, ProcessConfig, TriggerConfig


async def main() -> None:
    process = AgentProcess(
        name="triage",
        config=ProcessConfig(
            model="openai:gpt-5-mini",
            instructions="Triage new support tickets, then post one short summary.",
            servers={"tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"])},
            triggers=[TriggerConfig(type="cron", cron_expression="*/15 * * * *")],
        ),
    )
    await process.start()  # builds the agent, connects its tools, starts the trigger
    print(process.status()["state"], "with", process.status()["trigger_count"], "trigger")
    try:
        await asyncio.Event().wait()  # from here on, the agent wakes every 15 minutes
    finally:
        await process.stop()
        print(process.status()["state"])


if __name__ == "__main__":
    try:
        asyncio.run(main())
    except KeyboardInterrupt:
        pass
Output
running with 1 trigger
stopped

That's a working AI agent cron job. start() builds the agent, connects to its MCP tools and starts the trigger. From then on the cron trigger wakes it every 15 minutes, with no ainvoke() call from you, until you stop it. Everything else in this guide is what you add before you trust it to run unattended.


[02]

How a long-running agent works

An AgentProcess wraps an ordinary Promptise agent. Triggers put events on a queue. A worker takes each event, adds what the agent should know (its context state, recent conversation, remaining budget) and runs the agent once. After the run, the runtime checks the budget and decides whether the process keeps going.

Rendering diagram…

Each process moves through a small state machine, and every transition is validated:

State

How the process gets there

running

start(), or resume() after a pause

suspended

suspend(), a budget with on_exceeded="pause", or a health anomaly with on_anomaly="pause"

stopped

stop(), or a budget with on_exceeded="stop"

failed

A startup error, or max_consecutive_failures failed runs in a row (default 3)

A suspended process keeps its triggers. Events that fire while it's paused wait in the queue until someone calls resume().


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry: pip install promptise. Cron, webhook and file-watch triggers work out of the box; their dependencies ship with the base install.

  • An API key for a model provider. This guide uses OpenAI with openai:gpt-5-mini; set OPENAI_API_KEY in your environment or a .env file.


[04]

Build a ticket triage agent, step by step

The agent's job: every time it wakes, find new support tickets, give each a priority, and post one summary to the team channel, naming the on-call engineer for anything urgent.

Step 01

Give the agent its tools

The agent reaches your systems through MCP tools. Here a small MCP server stands in for a ticketing system and a team chat: a JSON file of tickets, and a log file as the channel. If you haven't built one before, How to Connect MCP Servers to Your AI Agent in Python covers the basics.

Pythontickets_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
import json
from datetime import datetime
from pathlib import Path
from typing import Literal

from promptise.mcp.server import MCPServer, ToolError

server = MCPServer("tickets")

# A stand-in for your ticketing system: a JSON file the agent reads and updates.
TICKETS = Path(__file__).with_name("tickets.json")
CHANNEL = Path(__file__).with_name("support-channel.log")


def load() -> dict:
    return json.loads(TICKETS.read_text()) if TICKETS.exists() else {}


@server.tool(read_only_hint=True)
async def list_untriaged() -> list[dict]:
    """List new support tickets that don't have a priority yet."""
    return [{"ticket_id": tid, **t} for tid, t in load().items() if t.get("priority") is None]


@server.tool()
async def set_priority(ticket_id: str, priority: Literal["P1", "P2", "P3"], reason: str) -> dict:
    """Give a ticket a priority. P1 = outage or data loss, P2 = a customer is blocked, P3 = everything else.

    Args:
        reason: One short sentence explaining the priority.
    """
    tickets = load()
    if ticket_id not in tickets:
        raise ToolError(f"There is no ticket {ticket_id}.")
    tickets[ticket_id].update(priority=priority, reason=reason)
    TICKETS.write_text(json.dumps(tickets, indent=2))
    return {"ticket_id": ticket_id, "priority": priority}


@server.tool()
async def post_summary(text: str) -> str:
    """Post a short triage summary to the support team's channel."""
    stamp = datetime.now().strftime("%H:%M:%S")
    with CHANNEL.open("a") as f:
        f.write(f"[{stamp}] {text}\n")
    return "posted"


if __name__ == "__main__":
    server.run()

Notice where the durable state lives: in the ticket system. A ticket with a priority is done, so the next run won't touch it again. That's the most reliable memory an agent can have, and it's yours, not the runtime's.

Step 02

Describe the process

ProcessConfig holds everything about the process: model, instructions, tools, triggers, starting state and limits. The file also writes a journal, which Step 6 explains.

Pythontriage_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
import asyncio
import sys

from promptise import CallbackSink, EventNotifier
from promptise.config import StdioServerSpec
from promptise.runtime import (
    AgentProcess,
    BudgetConfig,
    ContextConfig,
    FileJournal,
    JournalEntry,
    ProcessConfig,
    ToolCostAnnotation,
    TriggerConfig,
)

INSTRUCTIONS = """\
You triage the support queue. Each time you wake up:
1. Call list_untriaged.
2. Give every ticket a priority with set_priority.
3. If you set any priorities, post ONE short summary with post_summary.
   For every P1, name the on-call engineer: use the exact on_call value
   from [Context State]. Never invent a name.
If there are no new tickets, don't post anything and reply "Nothing to triage."
"""

CONFIG = ProcessConfig(
    model="openai:gpt-5-mini",
    instructions=INSTRUCTIONS,
    servers={
        "tickets": StdioServerSpec(command=sys.executable, args=["tickets_server.py"]),
    },
    # Every minute, for the demo. In production: "*/15 * * * *" or "0 8 * * 1-5".
    triggers=[TriggerConfig(type="cron", cron_expression="* * * * *")],
    # State the agent sees at the start of every run.
    context={"initial_state": {"on_call": "Sam"}},
    # The autonomy budget: what one run, and one day, may do.
    budget=BudgetConfig(
        enabled=True,
        max_tool_calls_per_run=6,
        max_runs_per_day=200,
        tool_costs={"post_summary": ToolCostAnnotation(cost_weight=1.0, irreversible=True)},
        max_irreversible_per_run=1,
        on_exceeded="pause",
    ),
)

journal = FileJournal(".promptise/journal")


def build_process(name: str = "triage", initial_state: dict | None = None) -> AgentProcess:
    config = CONFIG
    if initial_state is not None:
        config = CONFIG.model_copy(update={"context": ContextConfig(initial_state=initial_state)})

    async def record(event) -> None:
        """Write every runtime event to the journal, plus a checkpoint after each run."""
        await journal.append(JournalEntry(process_id=name, entry_type=event.event_type, data=event.data))
        if event.event_type == "invocation.complete":
            await journal.checkpoint(
                name,
                {
                    "context_state": process.context.state_snapshot(),
                    "lifecycle_state": process.state.value,
                    "budget_remaining": process.status().get("budget"),
                },
            )

    process = AgentProcess(
        name=name,
        process_id=name,
        config=config,
        event_notifier=EventNotifier(sinks=[CallbackSink(record)]),
    )
    return process


async def main() -> None:
    process = build_process()
    await process.start()
    print(f"{process.name} is {process.state.value}. Press Ctrl+C to stop.")
    try:
        while True:
            await asyncio.sleep(3600)
    finally:
        await process.stop()
        print(f"{process.name} is {process.state.value}.")


if __name__ == "__main__":
    try:
        asyncio.run(main())
    except KeyboardInterrupt:
        pass

The parts that matter:

  • The trigger. A standard five-field cron expression. This guide uses every minute so you can watch several cycles; a real triage agent might run every 15 minutes, or at 8:00 on weekdays.

  • The context. initial_state is the process's key-value state. The runtime shows it to the agent at the start of every run, and your code can change it while the process runs.

  • The budget. At most six tool calls per run, 200 runs a day, and one irreversible action per run. post_summary is marked irreversible because a message, once posted, can't be unposted. Cost weights are abstract units you choose, not money.

python triage_agent.py runs the process until you press Ctrl+C:

Output
triage is running. Press Ctrl+C to stop.
triage is stopped.
Step 03

Run it for a few cycles

To see the whole lifecycle in a few minutes, this script starts the process, adds tickets between runs and prints what happens. A small helper, seed_tickets.py, adds the tickets: two at first, one between runs, then a burst of five.

Pythondemo.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
"""Drive the triage process through three cron cycles and print what happens."""

import asyncio
import subprocess
import sys
import time
from pathlib import Path

from triage_agent import build_process

CHANNEL = Path("support-channel.log")


def seed(batch: str) -> None:
    out = subprocess.run([sys.executable, "seed_tickets.py", batch], capture_output=True, text=True)
    print(f"{time.strftime('%H:%M:%S')}  queue: {out.stdout.strip()}")


async def wait_for_run(process, n: int) -> None:
    while process.status()["invocation_count"] < n:
        await asyncio.sleep(1)
    await asyncio.sleep(1)  # let post-run bookkeeping finish
    s = process.status()
    print(
        f"{time.strftime('%H:%M:%S')}  run {n} done: state={s['state']} "
        f"conversation_messages={s['conversation_messages']} "
        f"tool_calls_left_this_run={s['budget']['tool_calls_run']}"
    )


async def main() -> None:
    CHANNEL.unlink(missing_ok=True)
    seed("1")
    process = build_process()
    await process.start()
    print(f"{time.strftime('%H:%M:%S')}  {process.name} is {process.state.value}")

    await wait_for_run(process, 1)

    # Shift change: tell the running agent who is on call now.
    process.context.put("on_call", "Dana", source="operator")
    seed("2")
    await wait_for_run(process, 2)

    # A burst of tickets: more work than one run's budget allows.
    seed("burst")
    await wait_for_run(process, 3)

    # Wait for the next cron tick while paused.
    await asyncio.sleep(65)
    s = process.status()
    print(f"{time.strftime('%H:%M:%S')}  after next tick: state={s['state']} queue_size={s['queue_size']}")

    print("\nsupport-channel.log:")
    print(CHANNEL.read_text())
    print("on_call history:", [(e.value, e.source) for e in process.context.state_history("on_call")])

    await process.stop()
    print(f"{time.strftime('%H:%M:%S')}  {process.name} is {process.state.value}")


asyncio.run(main())
Output
19:06:33  queue: 2 new ticket(s) in the queue
19:06:56  triage is running
19:07:12  run 1 done: state=running conversation_messages=2 tool_calls_left_this_run=2
19:07:12  queue: 1 new ticket(s) in the queue
19:08:09  run 2 done: state=running conversation_messages=4 tool_calls_left_this_run=3
19:08:09  queue: 5 new ticket(s) in the queue
Budget violation: max_tool_calls_per_run (limit=6, current=7) — action=pause
19:09:16  run 3 done: state=suspended conversation_messages=6 tool_calls_left_this_run=-1
19:10:21  after next tick: state=suspended queue_size=1

support-channel.log:
[19:07:09] P1: T-201 — Checkout 500s since 09:10 (on-call: Sam)
P2: T-202 — Customer can't download invoice PDF, blocking billing.
[19:08:06] P1: T-203 — Nobody in the EU region can log in. On-call: Dana.
[19:09:10] Triage: set priorities — T-300: P2 (password reset email not received); T-301: P2 (mobile app crashes opening settings); T-302: P2 (double charge); T-303: P3 (dark mode logo invisible); T-304: P2 (enterprise customer receiving 429). No P1s.

on_call history: [('Sam', 'system'), ('Dana', 'operator')]
19:10:21  triage is stopped

Three runs, each at the top of a minute, with no call from the script. Starting the process took about 20 seconds here; start() builds the agent and connects its tools once, and every run reuses them. The next three steps read this output from top to bottom.

Step 04

Change what the agent knows while it runs

Between runs 1 and 2, the on-call engineer changed. The script told the running process with one line:

Pythondemo.py
process.context.put("on_call", "Dana", source="operator")

Run 1 named Sam for its P1. Run 2 named Dana. Nothing restarted. The context keeps an audit trail too: state_history("on_call") shows who set each value, the initial system value and the operator change.

The process carries two kinds of memory between runs:

  • Context state, the key-value store you just used. The runtime adds it to the start of every run as a [Context State] message.

  • Recent conversation, a rolling buffer of each run's trigger message and the agent's final reply. That's why conversation_messages grows by two per run. The default keeps the last 100 messages; change it with conversation_max_messages in ContextConfig.

Warning

In 1.2.1, a long run can lose sight of its context state. Once a run has six tool results, the agent compacts its working context to the instructions, the latest tool calls and a list of facts gathered so far, and the [Context State] message is no longer in it. In a separate test with five new tickets, the model call that wrote the summary saw only that compacted view, and the summary named "alice@example.com" as on call. Keep runs short, and put facts the agent needs late in a run into its instructions or into a tool result.

Here's what the model received on each call of that test, captured with a LangChain callback:

Output
  MODEL SEES: system | You triage the support queue. Each time you wake up: 1. Call list_untriaged. 2. Give every ticket a priorit…
  MODEL SEES: system | [Context State] {'on_call': 'Dana'}
  MODEL SEES: human | [Trigger: cron] Payload: {}
  ---
…
  ---
  MODEL SEES: system | You triage the support queue. Each time you wake up: 1. Call list_untriaged. 2. Give every ticket a priorit…
  MODEL SEES: ai | 
  MODEL SEES: tool | {"ticket_id": "T-300", "priority": "P2"}
…
  MODEL SEES: system | Facts already gathered (do NOT call a tool to fetch any of these again): - list_untriaged({'payload': {}}) …
Step 05

Let the budget stop a runaway run

Run 3 got five tickets at once. Listing them, prioritizing all five and posting the summary takes seven tool calls; the budget allows six. Here's what happened, from the output above:

  • The run finished. All five tickets got a priority and the summary was posted.

  • After the run, the runtime saw the seventh call over the limit, logged Budget violation: max_tool_calls_per_run (limit=6, current=7) — action=pause, and suspended the process.

  • The next cron tick still fired. Its event waited in the queue (queue_size=1) instead of running.

That's the most important thing to know about the autonomy budget in 1.2.1: per-run limits are checked when the run ends, not in the middle of it. The run that crosses the line completes; the budget stops the next one. It's a circuit breaker for runaway schedules, not a hard cap on a single run. The daily limit, max_runs_per_day, is the exception: it's checked before a run starts, so the run over the limit never happens.

So set per-run limits a little above what a normal run needs, and use the daily limits to cap the total. Three more things worth knowing:

  • on_exceeded can be "pause", "stop" or "escalate". "escalate" posts the violation to a webhook or an event bus, then pauses.

  • Nothing resumes a paused process automatically, not even the daily reset. A person, or your own code, calls process.resume(), and the queued events run then.

  • With inject_remaining=True (the default), the agent sees its remaining budget at the start of each run as a [Budget Remaining] message.

Note

The budget counts tool calls and LLM turns, not money. Token spend isn't tracked. If you need a dollar limit, cap max_llm_turns_per_run and max_runs_per_day, and watch your provider's usage dashboard. The Autonomy Budget docs cover cost weights in detail.

Step 06

Keep a journal you can read back

When something goes wrong at 3 a.m., you want a record of what the agent did. Promptise has a journal for exactly this: FileJournal writes append-only JSONL, one file per process, and promptise runtime logs prints it.

Important

In 1.2.1, AgentProcess doesn't write the journal itself. ProcessConfig accepts a journal=JournalConfig(...) setting, but a test run with JournalConfig(level="full") created no journal file at all. Until that's wired up, write the entries yourself, as below.

The process emits events as it works: process.started, invocation.start, invocation.complete, budget.warning, budget.exceeded and more. Pass an EventNotifier with a CallbackSink, and you can journal every one of them. This is the record function from Step 2:

Pythontriage_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
    async def record(event) -> None:
        """Write every runtime event to the journal, plus a checkpoint after each run."""
        await journal.append(JournalEntry(process_id=name, entry_type=event.event_type, data=event.data))
        if event.event_type == "invocation.complete":
            await journal.checkpoint(
                name,
                {
                    "context_state": process.context.state_snapshot(),
                    "lifecycle_state": process.state.value,
                    "budget_remaining": process.status().get("budget"),
                },
            )

After each run it also stores a checkpoint: a snapshot of the context state, the lifecycle state and the remaining budget. Read the journal back from the command line:

>_Terminal
promptise runtime logs triage --lines 40
Output
                                Journal: triage                                 
┏━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Time                    ┃ Type                 ┃ Data                        ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ 2026-10-10T17:06:56.963 │ process.started      │ {"process_name": "triage",  │
│                         │                      │ "process_id": "triage"}     │
│ 2026-10-10T17:07:00.118 │ invocation.start     │ {"model":                   │
│                         │                      │ "openai:gpt-5-mini"}        │
│ 2026-10-10T17:07:10.289 │ invocation.complete  │ {"duration_ms": 10223.2}    │
│ 2026-10-10T17:07:10.290 │ checkpoint           │ {"state": {"context_state": │
│                         │                      │ {"on_call": "Sam"},         │
│                         │                      │ "lifecycle_state":          │
│                         │                      │ "running",...               │
…
│ 2026-10-10T17:09:14.947 │ invocation.complete  │ {"duration_ms": 14936.9}    │
│ 2026-10-10T17:09:14.948 │ checkpoint           │ {"state": {"context_state": │
│                         │                      │ {"on_call": "Dana"},        │
│                         │                      │ "lifecycle_state":          │
│                         │                      │ "suspende...                │
│ 2026-10-10T17:09:14.956 │ budget.exceeded      │ {"process_name": "triage",  │
│                         │                      │ "limit_type":               │
│                         │                      │ "max_tool_calls_per_run",   │
│                         │                      │ "current":...               │
└─────────────────────────┴──────────────────────┴─────────────────────────────┘

The full entries are in .promptise/journal/triage.jsonl, for example this line from run 3:

JSON
{"entry_id": "…", "process_id": "triage", "timestamp": "2026-10-10T17:09:14.956…", "entry_type": "budget.exceeded", "data": {"process_name": "triage", "limit_type": "max_tool_calls_per_run", "current": 7, "limit": 6}}

promptise runtime logs looks up the journal by the name you give it, which is why build_process uses the process name as its process_id too.

Step 07

Restart without losing state

The process keeps its context in memory, so a restart starts from the initial state again. The checkpoint fixes that. ReplayEngine reads the journal and returns the last known state, and you start the new process with it:

Pythonrecover.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
import asyncio

from promptise.runtime import ReplayEngine

from triage_agent import build_process, journal


async def main() -> None:
    recovered = await ReplayEngine(journal).recover("triage")
    print("recovered:", recovered)

    process = build_process(initial_state=recovered["context_state"])
    await process.start()
    print("state after restart:", process.context.state_snapshot())
    await process.stop()


asyncio.run(main())
Output
recovered: {'context_state': {'on_call': 'Dana'}, 'lifecycle_state': 'suspended', 'last_entry_type': 'budget.warning', 'entries_replayed': 10}
state after restart: {'on_call': 'Dana'}

The new process knows Dana is on call, although the config still says Sam. It also learned the old process was suspended when it stopped: here, paused by its budget. Starting it again is a decision for a person, so check lifecycle_state before you do.

Warning

In 1.2.1, ReplayEngine replays journal entries from the first checkpoint, not the last. If you also journal context_update entries, an old value can overwrite the newer checkpoint. Journal checkpoints and events, as above, and the recovered state is the latest checkpoint.


[05]

Wake on a webhook instead of a clock

A schedule is right for sweeps. When something else already knows a ticket arrived, let it wake the agent. A webhook trigger starts a small HTTP server inside the process; this keeps a slow cron sweep as a safety net:

Pythonwebhook_agent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
import asyncio

from promptise.runtime import AgentProcess, TriggerConfig

from triage_agent import CONFIG

config = CONFIG.model_copy(
    update={
        "triggers": [
            TriggerConfig(type="cron", cron_expression="*/15 * * * *"),
            TriggerConfig(type="webhook", webhook_path="/tickets", webhook_port=8180),
        ]
    }
)


async def main() -> None:
    process = AgentProcess(name="triage", config=config)
    await process.start()
    print("listening on http://127.0.0.1:8180/tickets")
    while process.status()["invocation_count"] < 1:
        await asyncio.sleep(1)
    print("runs:", process.status()["invocation_count"])
    await process.stop()


asyncio.run(main())

Post an event the way your help desk would:

>_Terminal
curl -s -X POST http://127.0.0.1:8180/tickets -H 'Content-Type: application/json' -d '{"event": "ticket.created", "ticket_id": "T-401"}'
Output
{"status": "accepted", "event_id": "494fb90b-8bd7-4f7e-a816-bef0093f1e0e"}

The webhook answers 202 at once and the run happens in the background. The JSON body becomes the trigger payload the agent sees. The process printed:

Output
WebhookTrigger on port 8180 has no HMAC secret — any HTTP client can trigger this webhook. Set hmac_secret for production use.
listening on http://127.0.0.1:8180/tickets
runs: 1

And the channel got:

Output
[19:13:16] P1: T-401 — Data export deletes rows from source table (possible data loss). On-call: Sam.

Take that warning seriously. A webhook created from TriggerConfig listens on 127.0.0.1 only and has no signature check, and TriggerConfig has no field to change either. The WebhookTrigger class itself takes host and hmac_secret, so register a trigger type that builds it with a secret:

Pythonprobe_trigger_config_gaps.py
register_trigger_type(
    "signed_webhook",
    lambda config, **_: WebhookTrigger(path=config.webhook_path, port=config.webhook_port,
                                       hmac_secret=config.custom_config["secret"]),
)

Use it as TriggerConfig(type="signed_webhook", ..., custom_config={"secret": ...}). In a test, an unsigned POST got 401 and a POST with a valid X-Webhook-Signature: sha256=<hex> header got 202.

The other built-in triggers work the same way. file_watch wakes the agent when files matching watch_patterns appear or change in watch_path. event and message listen to an event bus or message broker, which you pass to an AgentRuntime to share between its processes. Triggers lists them all, and how to register your own.


[06]

Declare it in a manifest and run it from the CLI

You don't have to write Python to run a process. A .agent manifest describes the same thing in YAML, and the promptise runtime command runs it:

YAMLtriage.agent
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
version: "1.0"
name: triage
model: openai:gpt-5-mini
instructions: |
  You triage the support queue. Each time you wake up:
  1. Call list_untriaged.
  2. Give every ticket a priority with set_priority.
  3. If you set any priorities, post ONE short summary with post_summary.
     For every P1, name the on-call engineer: use the exact on_call value
     from [Context State]. Never invent a name.
  If there are no new tickets, don't post anything and reply "Nothing to triage."

servers:
  tickets:
    type: stdio
    command: python
    args: ["tickets_server.py"]

triggers:
  - type: cron
    cron_expression: "* * * * *"   # every minute for the demo

world:
  on_call: Sam

budget:
  enabled: true
  max_tool_calls_per_run: 6
  max_runs_per_day: 200
  max_irreversible_per_run: 1
  on_exceeded: pause
  tool_costs:
    post_summary:
      cost_weight: 1.0
      irreversible: true

world is the manifest's name for the initial context state. Check the file before you run it:

>_Terminal
promptise runtime validate triage.agent
Output
Validating triage.agent...
✓ Schema validation passed
✓ No warnings
                           Manifest Summary                            
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Field        ┃ Value                                                ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ Name         │ triage                                               │
│ Model        │ openai:gpt-5-mini                                    │
│ Version      │ 1.0                                                  │
│ Instructions │ You triage the support queue. Each time you wake up: │
│              │ 1. Call...                                           │
│ Servers      │ 1                                                    │
│ Triggers     │ 1                                                    │
└──────────────┴──────────────────────────────────────────────────────┘
✓ Validation complete!

Then start it. It runs in the foreground until you press Ctrl+C:

>_Terminal
promptise runtime start triage.agent
Output
Loaded process: triage
                        Agent Processes                         
┏━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┓
┃ Name   ┃ State   ┃ PID      ┃ Invocations ┃ Queue ┃ Uptime   ┃
┡━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━┩
│ triage │ RUNNING │ 51495264 │           0 │     0 │ 0h 0m 0s │
└────────┴─────────┴──────────┴─────────────┴───────┴──────────┘

Press Ctrl+C to stop all processes.

Shutting down...
All processes stopped.

In that run the agent woke at the next minute, triaged two tickets and posted its summary before Ctrl+C. Pass a directory instead of a file to start every .agent file in it, and promptise runtime init --template cron writes a starter manifest. The Agent Manifests and Runtime CLI docs have every field and option.

Note

In 1.2.1, promptise runtime stop, status and restart only print a message; they can't reach a running process. And --detach prints "Running in background." but nothing is left running once the command returns. Run promptise runtime start in the foreground under a supervisor such as systemd, Docker or Kubernetes, and let that handle stopping and restarting.


[07]

Other guardrails: health, mission and secrets

The budget limits how much an agent does. Three more opt-in subsystems watch what it does. All are off by default.

  • Behavioral health (HealthConfig) spots bad patterns without calling a model: the same tool with the same arguments again and again, repeating loops, empty responses and high error rates. on_anomaly can log, pause or escalate.

  • Mission (MissionConfig) gives a process an objective and success_criteria. Every eval_every runs, a model judges progress; the process can stop itself when the mission is achieved, or escalate when confidence drops below confidence_threshold.

  • Secret scoping (SecretScopeConfig) gives each process its own secrets, resolved from ${ENV_VAR} references when it starts, encrypted in memory, with optional expiry, and revoked on stop(). A missing variable stops the start with KeyError: "Environment variable 'NOT_SET_ANYWHERE' not found (required by secret config)". The agent itself can only read them through the get_secret tool in open execution mode.

Tune health checks before you let them pause anything. With the default thresholds, an agent whose queue was empty three runs in a row was paused as unhealthy, because three short tool results in a row ([]) count as empty responses:

Output
run 1: state=running anomalies=0
run 2: state=running anomalies=0
Behavioral anomaly detected: empty_response — Agent producing empty responses: 3 consecutive responses under 10 characters
run 3: state=suspended anomalies=1

Start with on_anomaly="log", read what it flags for a few days, then raise empty_threshold and stuck_threshold to fit your agent before switching to "pause". See Behavioral Health, Mission Model and Secret Scoping.


[08]

Before you go to production

Know where every piece of state lives. In 1.2.1:

State

Where it lives

Survives a restart?

Context state

Process memory

Only if you checkpoint and restore it, as in Step 7

Recent conversation

Process memory; stop() clears it

No

Budget counters, daily ones too

Process memory

No: they start from zero

Journal and checkpoints

Files you write

Yes

Tickets, orders, records

Your systems, behind MCP tools

Yes

With that in mind:

  • Restarts are your supervisor's job. ProcessConfig has restart_policy and max_restarts, but 1.2.1 doesn't act on them. In a test with restart_policy="on_failure", a process that failed three runs in a row went to failed and was still failed 15 seconds later. Run each process under systemd, Docker or Kubernetes with a restart policy, restore context from your checkpoint on start, and alert on failed.

  • Run one copy of each process. Each process has its own cron trigger and nothing coordinates replicas, so two copies of the same process both fire at every tick and do the work twice. Scale by giving different processes different jobs, and make tools that change data safe to repeat.

  • Distributed means remote control, not failover. RuntimeTransport puts an HTTP API on a node's runtime (health, status, start, stop, inject an event), with an optional bearer token. RuntimeCoordinator registers nodes, checks their health and aggregates status. There's no shared state, no leader election and no automatic move of processes off a failed node. Also, the coordinator can't send a token: against a node with auth_token set, its health check still worked, but get_node_status failed with Status request failed: 401.

  • Keep runs small. Short runs stay inside the budget, keep the context state in view (see Step 4), and are cheaper to retry.

  • Lock down triggers. Webhooks need a signature check and a reverse proxy in front if anything outside the machine calls them.

  • Watch the journal. Alert on budget.exceeded, health.anomaly and invocation.error events; the CallbackSink from Step 6 is a good place to send them on.


[09]

Frequently asked questions

How do I run an AI agent as a cron job?

Give an AgentProcess a trigger like TriggerConfig(type="cron", cron_expression="*/15 * * * *") and start it. The process stays up and wakes the agent on schedule, so the agent and its tool connections are built once instead of on every run. A plain system cron job that runs a script also works; it just rebuilds everything each time and keeps no state in between.

How does a long-running AI agent remember things between runs?

Three ways, in order of reliability. Your own systems, through tools: a triaged ticket stays triaged. The process context, a key-value store the agent sees at the start of every run, which you can checkpoint to a journal and restore. And a rolling buffer of recent trigger messages and replies. For anything that must survive a restart, use the first two.

How do I stop an autonomous AI agent from running up costs?

Enable BudgetConfig with per-run and daily limits on tool calls, LLM turns, cost weights and irreversible actions, and choose pause, stop or escalate for when a limit is crossed. Per-run limits take effect after the run that crosses them; max_runs_per_day is checked before a run starts, so it's your hard cap. Track real token spend with your model provider as well.

Can I run AI agents on several machines?

You can run a runtime on each machine and manage them over HTTP with RuntimeTransport and RuntimeCoordinator. In 1.2.1 that gives you remote start, stop, status and health checks, not shared state or automatic failover, so assign each process to one machine.

Do I need Kubernetes to deploy an AI agent?

No. One promptise runtime start in the foreground under any supervisor that restarts it on failure is enough to begin with: systemd, Docker with a restart policy, or a Kubernetes Deployment with one replica.


[10]

Where to go next

  • Agent Runtime: the overview of processes, triggers, state and governance.

  • Agent Processes: every ProcessConfig field and lifecycle method.

  • Cron triggers and Triggers: schedules, webhooks, file watches and custom trigger types.

  • Journal and Context & State: what to record and what the agent sees.

  • Runtime Manager: several processes in one AgentRuntime.

  • Build a Production MCP Server in Python: make the tools this agent calls safe for unattended use.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: give the agent tools for an API you already have.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths