An agent defined in Python lives inside someone's script. To change its model or its instructions, you need that person, a code review of a Python file, and a redeploy. Defining the agent in a YAML file fixes that: the agent becomes a short config file that sits in Git, gets reviewed like any other change, and runs from the command line without code. In this guide you'll write an on-call assistant for a platform team as a YAML file, validate it, run it from the CLI and from Python, keep its secrets out of the file, and set up a check that catches broken agent files before they merge. Every snippet was run against Promptise Foundry 1.2.1, and every output is what it printed.
How do you define an AI agent in a YAML file?
You write the model, the instructions and the MCP servers the agent may use into one file, then point a loader at it. In Promptise Foundry that file is a SuperAgent file, a plain YAML document with the .superagent extension:
version: "1.0"
agent:
model: "openai:gpt-5-mini"
instructions: |
You are the on-call assistant for the platform team.
Use your tools to answer questions about incidents and runbooks.
Quote runbook steps exactly. Never invent incidents.
trace: true
servers:
incidents:
type: stdio
command: python
args: ["incidents_server.py"]
env:
PAGER_API_TOKEN: "${PAGER_API_TOKEN}"Check it, then run it:
promptise validate oncall.superagent
promptise agent oncall.superagentThat's the whole agent. There's no Python in it, ${PAGER_API_TOKEN} is filled in from the environment when the file loads, and a misspelled key is rejected instead of ignored. The same file also loads from Python with load_superagent_file, so the CLI and your services run exactly the same definition.
[02]
How it works
The file is a contract. Promptise reads it, checks it against a strict schema, fills in the environment variables, and turns it into the same arguments you'd otherwise pass to build_agent by hand:
Rendering diagram…
Every section in the schema forbids unknown keys, so instuctions is an error rather than an agent that quietly runs without its instructions. These are the top-level sections a SuperAgent file can have in 1.2.1:
Section | What it sets |
|---|---|
version | The schema version. Only "1.0" exists, and it's the default. |
agent | Required. The model, the instructions and trace, which prints every tool call. |
servers | Named MCP servers, each type: stdio or type: http. |
cross_agents | Other .superagent files this agent can delegate to. |
identity | Who the agent is: a local name for attribution, or a verifiable identity from Entra, AWS, GCP, SPIFFE or OIDC. |
memory | Long-term memory: in_memory, chroma or mem0. |
sandbox | A container the agent can run code in. |
approval | Tools that need a person's approval before they run. |
guardrails, cache, observability, optimize_tools, events, adaptive | Guardrails, semantic caching, observability, tool optimization, event notifications and adaptive strategy, the same features build_agent offers. |
max_invocation_time | A time limit per request, in seconds. 0, the default, means no limit. |
A file needs at least one of servers, cross_agents or sandbox, because an agent with nothing to work with isn't much use. The SuperAgent files reference lists every field in every section.
[03]
What you need
Python 3.10 or newer.
Promptise Foundry from PyPI. It includes the promptise command.
An API key for a model provider. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.
pip install promptise
export OPENAI_API_KEY="sk-..."[04]
Build a YAML-defined agent, step by step
The agent answers on-call questions: what's broken, and what the runbook says to do about it. A small MCP server gives it the data.
Write the MCP server
The tools live in an MCP server, not in the YAML file. That's what keeps the file short: it says which tools the agent may use, and the server decides what they do.
import os
import sys
from promptise.mcp.server import MCPServer
# The token this server would use to call your paging system.
# It comes from the environment, never from code or config files.
if not os.environ.get("PAGER_API_TOKEN"):
sys.exit("PAGER_API_TOKEN is not set")
server = MCPServer("incidents")
# A stand-in for your incident tracker.
INCIDENTS = [
{"id": "INC-311", "service": "checkout", "severity": "high", "status": "open",
"summary": "Payment confirmations delayed by up to 4 minutes"},
{"id": "INC-309", "service": "search", "severity": "low", "status": "open",
"summary": "Autocomplete returns stale results for new products"},
{"id": "INC-302", "service": "checkout", "severity": "medium", "status": "resolved",
"summary": "Coupon codes rejected for EU customers"},
]
RUNBOOKS = {
"checkout": "1. Check the payments queue depth. 2. If above 5,000, scale the "
"payment-workers deployment to 12 replicas. 3. Page #payments if it keeps growing.",
"search": "1. Check the last index build. 2. If older than 1 hour, trigger a reindex.",
}
@server.tool()
async def list_incidents(status: str = "open") -> list[dict]:
"""List incidents by status.
Args:
status: "open" or "resolved". Defaults to "open".
"""
return [i for i in INCIDENTS if i["status"] == status]
@server.tool()
async def get_runbook(service: str) -> str:
"""Get the on-call runbook for a service, for example "checkout".
Args:
service: The service name.
"""
return RUNBOOKS.get(service.strip().lower(), f"No runbook for {service}.")
if __name__ == "__main__":
server.run()The server refuses to start without PAGER_API_TOKEN, like a real integration would. You'll see why that matters in a moment. If MCP servers are new to you, How to Connect MCP Servers to Your AI Agent in Python walks through one line by line.
Describe the agent in YAML
Save the file from the top of this guide as oncall.superagent, next to the server. Here's what each part does:
`agent.model` is the same provider:model string build_agent takes. For a temperature, a token limit or an Azure deployment, model also accepts a mapping with provider, name and those settings.
`agent.instructions` is the system prompt. YAML's | keeps the line breaks, so a long prompt stays readable, and every edit shows up as a clean line in a diff.
`servers.incidents` starts incidents_server.py as a child process over stdio. command: python is looked up on your PATH, so run the agent with your virtual environment active.
`env` passes PAGER_API_TOKEN to that child process. The value is a placeholder; the real token never appears in the file.
Validate the file
promptise validate checks the file without starting anything: the YAML syntax, the schema, and whether every variable the file mentions is set.
promptise validate oncall.superagentValidating oncall.superagent...
✓ File format and schema valid
✓ All environment variables available
✓ Validation complete!Run it in a shell where PAGER_API_TOKEN isn't set, and it tells you:
Validating oncall.superagent...
✓ File format and schema valid
⚠ Missing environment variables:
- PAGER_API_TOKEN
Note: Set these variables or use defaults (${VAR:-default}) to resolve.
✓ Validation complete!Notice that this is a warning, and the command still exits with status 0. The agent itself is stricter: promptise agent refuses to start until the variable is set.
Failed to load configuration: Missing required environment variables in …/oncall-agent/oncall.superagent: PAGER_API_TOKENRun it from the CLI
promptise agent loads the file, builds the agent and opens a chat in your terminal. Type exit to leave.
export PAGER_API_TOKEN="..."
promptise agent oncall.superagentPromptise Foundry loaded from oncall.superagent. Type 'exit' to quit.
> → Invoking tool: list_incidents with {'status': 'open'}
→ Invoking tool: get_runbook with {'service': 'checkout'}
✔ Tool result from list_incidents: [{"id": "INC-311", "service": "checkout", "severity": "high", "status": "open", "summary": "Payment confirmations delayed by up to 4 minutes"}, {"id": "INC-309", "service": "search", "severity": "low", "status": "open", "summary": "Autocomplete returns stale results for new products"}]
✔ Tool result from get_runbook: 1. Check the payments queue depth. 2. If above 5,000, scale the payment-workers deployment to 12 replicas. 3. Page #payments if it keeps growing.
╭────────────────────────────── Final LLM Answer ──────────────────────────────╮
│ Yes. │
│ │
│ Open checkout incident: │
│ - INC-311 (service: checkout, severity: high) — Payment confirmations │
│ delayed by up to 4 minutes │
│ │
│ Runbook (quoted exactly): │
│ "1. Check the payments queue depth. 2. If above 5,000, scale the │
│ payment-workers deployment to 12 replicas. 3. Page #payments if it keeps │
│ growing." │
╰──────────────────────────────────────────────────────────────────────────────╯This run asked "Is anything open on checkout, and what does the runbook say to do?" at the > prompt; the question was piped in, so it isn't echoed above. The agent called both tools at once, picked the checkout incident out of the list, and quoted the runbook word for word, as its instructions asked. The only Python involved is the server; the agent itself is the 17-line file.
The CLI can also change a setting for one run without touching the file. --model-id, --instructions and --trace/--no-trace override the file, and --stdio or --http add a server next to the ones it declares. That's handy for trying a new model before you propose the change.
Run the same file from Python
Your services can load the exact file the CLI runs. load_superagent_file validates it and fills in the variables, and to_build_kwargs() hands the result to build_agent:
import asyncio
import sys
from promptise import build_agent
from promptise.exceptions import SuperAgentError, SuperAgentValidationError
from promptise.models import load_dotenv_if_present
from promptise.superagent import load_superagent_file
async def main(path: str, question: str) -> None:
# The CLI reads .env for you. In your own code, do it before loading the file.
load_dotenv_if_present()
try:
loader, _ = load_superagent_file(path)
except SuperAgentValidationError as exc:
for error in exc.errors:
print(" -> ".join(map(str, error["loc"])), ":", error["msg"])
raise SystemExit(1)
except SuperAgentError as exc:
raise SystemExit(str(exc))
config = loader.to_agent_config()
print("Model:", config.model)
print("Servers:", list(config.servers))
agent = await build_agent(**config.to_build_kwargs())
try:
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
print(">>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main(sys.argv[1], sys.argv[2]))python run_agent.py oncall.superagent "Which incidents are open right now, most severe first?"Model: openai:gpt-5-mini
Servers: ['incidents']
→ Invoking tool: list_incidents with {'status': 'open'}
✔ Tool result from list_incidents: [{"id": "INC-311", "service": "checkout", "severity": "high", "status": "open", "summary": "Payment confirmations delayed by up to 4 minutes"}, {"id": "INC-309", "service": "search", "severity": "low", "status": "open", "summary": "Autocomplete returns stale results for new products"}]
>>> Open incidents (most severe first):
1. INC-311 — service: checkout — severity: high — status: open
Summary: Payment confirmations delayed by up to 4 minutes
2. INC-309 — service: search — severity: low — status: open
Summary: Autocomplete returns stale results for new productsTwo details in the script are worth copying:
`load_dotenv_if_present()` loads the nearest .env file, the way the CLI does on every start. Without it, load_superagent_file only sees variables that are already exported, and a key that lives in .env counts as missing.
`SuperAgentValidationError` carries the problems as a list in exc.errors, each with a loc path and a msg. That's what you'd log, or show in an internal tool where people edit agent files.
load_superagent_file returns a second value too: the loaded files from cross_agents, which this agent doesn't use. If yours does, see the limits below.
Catch mistakes before they ship
Here's the same agent with two typos, the kind that slip through a quick edit:
version: "1.0"
agent:
model: "openai:gpt-5-mini"
instuctions: |
You are the on-call assistant for the platform team.
trace: true
servers:
incidents:
type: stdio
comand: python
args: ["incidents_server.py"]promptise validate broken.superagentValidating broken.superagent...
✗ Schema validation failed:
Schema validation failed for …/oncall-agent/broken.superagent
File: …/oncall-agent/broken.superagent
Validation errors:
agent -> instuctions: Extra inputs are not permitted
servers -> incidents -> stdio -> command: Field required
servers -> incidents -> stdio -> comand: Extra inputs are not permittedThis time the exit status is 1. Each line points at the exact key: the misspelled instuctions, and comand, which is both an unknown key and the reason command is missing. Loading the file from Python stops with the same errors, which run_agent.py prints from exc.errors:
agent -> instuctions : Extra inputs are not permitted
servers -> incidents -> stdio -> command : Field required
servers -> incidents -> stdio -> comand : Extra inputs are not permittedWithout the strict schema, the first typo would give you an agent that runs fine and ignores its instructions, which is much harder to notice than an error.
[05]
Keep secrets out of the file
The agent file belongs in Git, so nothing secret goes in it. Every string in the file can reference an environment variable instead:
You write | You get |
|---|---|
| The variable's value. Loading stops if it isn't set. |
| The variable's value, or the default after :- when it isn't set. |
| Placeholders work inside longer strings too. |
The variables come from wherever you already keep them: your shell, your CI secrets, your container platform or a secret manager that exports them. The Environment Resolver page covers the syntax in full.
Pass secrets to stdio servers explicitly
A stdio server doesn't see your whole environment. It inherits only a short list of safe variables, HOME, LOGNAME, PATH, SHELL, TERM and USER on macOS and Linux, plus whatever you list under env. Here's the same agent file with the env block removed, run in a shell where PAGER_API_TOKEN is exported:
PAGER_API_TOKEN is not set
…
MCPClientError: Failed to connect to server 'incidents': Failed to connect to python incidents_server.py (McpError: Connection closed)The first line is the server's own message: it never received the token. That's a good default. Each server gets only the secrets the file names, and a reviewer can see exactly which ones those are.
Use .env for local development
The CLI loads a .env file from the folder you run it in, or the nearest parent up to your project root, and never overwrites a variable that already has a value. In Python, load_dotenv_if_present() does the same, as in run_agent.py. Keep .env out of the repository:
.gitignore:1:.env .envThat's git check-ignore -v .env confirming the rule. Model Setup explains how .env and exported variables work together.
Let the schema warn you about pasted keys
If someone pastes a model key into the file, validation warns them. Here model uses the detailed form with an api_key:
Validating hardcoded-key.superagent...
…/pydantic/main.py:790: UserWarning: Direct API key detected in config. Consider using ${ENV_VAR} syntax for security.
return cls.__pydantic_validator__.validate_python(
✓ File format and schema valid
✓ All environment variables available
✓ Validation complete!It's only a warning, and it only checks model.api_key. A token pasted into headers or env gets no warning, so keep a secret scanner in CI and an eye on diffs.
[06]
Connect to a remote MCP server with an API key
When the incidents server runs as a shared service, the agent connects over HTTP instead of starting it. Here the server requires an API key on every call:
import os
from promptise.mcp.server import APIKeyAuth, AuthMiddleware, MCPServer
server = MCPServer("incidents", require_auth=True)
server.add_middleware(
AuthMiddleware(APIKeyAuth(keys={os.environ["INCIDENTS_API_KEY"]: "oncall-agent"}))
)
if __name__ == "__main__":
server.run(transport="http", host="127.0.0.1", port=8270)The tools are the same as before. The agent file sends the key in the x-api-key header, read from the environment, and takes the URL from a variable with a local default:
version: "1.0"
agent:
model: "openai:gpt-5-mini"
instructions: |
You are the on-call assistant for the platform team.
Use your tools to answer questions about incidents and runbooks.
Quote runbook steps exactly. Never invent incidents.
servers:
incidents:
type: http
url: "${INCIDENTS_URL:-http://127.0.0.1:8270/mcp}"
headers:
x-api-key: "${INCIDENTS_API_KEY}"python run_agent.py oncall-remote.superagent "What should I do about the checkout incident?"Model: openai:gpt-5-mini
Servers: ['incidents']
→ Invoking tool: list_incidents with {'status': 'open'}
→ Invoking tool: get_runbook with {'service': 'checkout'}
✔ Tool result from list_incidents: [{"id": "INC-311", "service": "checkout", "severity": "high", "status": "open", "summary": "Payment confirmations delayed by up to 4 minutes"}, {"id": "INC-309", "service": "search", "severity": "low", "status": "open", "summary": "Autocomplete returns stale results for new products"}]
✔ Tool result from get_runbook: 1. Check the payments queue depth. 2. If above 5,000, scale the payment-workers deployment to 12 replicas. 3. Page #payments if it keeps growing.
>>> The active checkout incident is INC-311 (severity: high) — "Payment confirmations delayed by up to 4 minutes."
…There's no trace: true in this file, and the tool calls still print: trace defaults to on. Set trace: false for agents that run unattended. For a server that expects a JSON Web Token, put Authorization: "Bearer ${API_TOKEN}" under headers instead. Authentication & Security covers the server side.
[07]
Review agent changes in Git
Once the agent is a file, changing it is a pull request. Here a teammate tightens the instructions and adds a time limit per request:
diff --git a/oncall.superagent b/oncall.superagent
index 4435c14..c69752b 100644
--- a/oncall.superagent
+++ b/oncall.superagent
@@ -6,6 +6,7 @@ agent:
You are the on-call assistant for the platform team.
Use your tools to answer questions about incidents and runbooks.
Quote runbook steps exactly. Never invent incidents.
+ Never suggest scaling a service without naming the incident it fixes.
trace: true
servers:
@@ -15,3 +16,5 @@ servers:
args: ["incidents_server.py"]
env:
PAGER_API_TOKEN: "${PAGER_API_TOKEN}"
+
+max_invocation_time: 60A reviewer sees the full behaviour change in two hunks, without reading any Python. The same goes for the changes that matter most: a new server, a different model, a new secret in an env block, or a removed approval rule each show up as a line in the diff.
Make the review safe to approve by validating every agent file on each pull request. This script finds them with git ls-files and stops at the first one that fails:
#!/usr/bin/env bash
# Validate every agent file in the repository. Run it in CI on each pull request.
set -euo pipefail
git ls-files '*.superagent' | while read -r file; do
promptise validate "$file" --no-check-env
done--no-check-env skips the variable check, because CI doesn't hold your production secrets and that check never fails the command anyway. On the change above, it passes:
Validating oncall-remote.superagent...
✓ File format and schema valid
✓ Validation complete!
Validating oncall.superagent...
✓ File format and schema valid
✓ Validation complete!On a branch that sets transport: websocket on the remote server, it fails with status 1 and names the field and the allowed values:
Validating oncall-remote.superagent...
✗ Schema validation failed:
…
Validation errors:
servers -> incidents -> http -> transport: Input should be 'http', 'streamable-http' or 'sse'To decide who may approve agent changes, put the agent files in their own folder and give it an owner with a CODEOWNERS file, so a platform lead signs off on every change to what the agents can reach.
[08]
Honest limits
These are true of Promptise Foundry 1.2.1, and worth knowing before you hand agent files to a team:
`validate` checks the schema, not whether the agent can start. It doesn't contact the model or start servers. Some settings also pass the schema and only fail at start-up: approval.handler: callback validates cleanly, and then promptise agent stops with
ValueError: Unsupported approval handler type: 'callback'. The same happens with handler: webhook and no webhook_url. Start each agent once in a staging environment before you rely on a change.Approval from YAML is limited. callback fails at start-up, as above. queue builds a queue that nothing outside the process can answer: from the CLI, a gated call waited out its timeout and the model was told
Approval timed out after 5.0s. webhook works, but the file has no field for its signing secret, so each process signs with a random one your approvals service can't verify. For approvals you can trust, build the agent in Python with an ApprovalPolicy, as in the Human-in-the-Loop Approval docs.Missing variables are only a warning in `validate`. It exits with status 0. Check variables where the agent runs: promptise agent and load_superagent_file both refuse to start without them.
Relative paths follow your working directory, not the file. Running promptise agent oncall-agent/oncall.superagent from the parent folder failed, because python looked for incidents_server.py in the parent folder. Run agents from the folder that holds the file, or set cwd on the server to an absolute path.
The CLI is an interactive chat. There's no single-question mode, and a server that fails to start prints a full traceback rather than one line. For jobs and services, load the file from Python as in Step 5.
Delegated agents don't get every section. The CLI builds each cross_agents file from its model, servers, instructions and trace only. In Python, to_build_kwargs() leaves cross_agents out entirely, so you build those agents yourself, as the Cross-Agent Delegation docs show.
[09]
Frequently asked questions
What is an AI agent configuration file?
A file that describes an agent instead of building it in code: which model it uses, its instructions, which tools or MCP servers it may call, and settings like memory or approval. A runtime reads the file and builds the agent from it. In Promptise Foundry that's a .superagent YAML file, run with promptise agent or loaded with load_superagent_file.
What are declarative AI agents?
Agents you describe rather than program. You say what the agent is and what it may use, and the framework turns that into a running agent. The behaviour still comes from the model and the tools; the declaration just moves the wiring out of code, so people who don't write Python can read, review and change it.
What are AI agents as code?
Treating agent definitions like infrastructure as code: plain text files in Git, changed through pull requests, checked automatically in CI and deployed from the repository. You get history, review and rollback for free, because a bad change is one git revert away.
Can I use a .yaml extension instead of .superagent?
Yes, as long as the name still says what it is: oncall.superagent.yaml and oncall.superagent.yml both load. A plain oncall.yaml is refused with Invalid file extension. Expected .superagent, .superagent.yaml, or .superagent.yml. The double extension also lets editors apply YAML highlighting.
Can I start a new agent file from a template?
Yes. promptise init writes a starter file, and --template picks one of basic, http, stdio, cross-agent or advanced. The templates use openai:gpt-4.1 as the model, so change that line to the model you use.
[10]
Where to go next
SuperAgent files: every section and field, with examples.
SuperAgent API reference: SuperAgentLoader, load_superagent_file and SuperAgentConfig.
CLI reference: promptise agent, validate and init with all their options.
Environment Resolver: the ${VAR} syntax and its helpers.
Server configuration: what the stdio and http server entries become in code.
How to Connect MCP Servers to Your AI Agent in Python: the same agent built in Python.
OpenAPI to MCP: Turn Any REST API into an MCP Server: generate a server your agent file can point at.