Most AI agents call their tools with a shared API key, or with whatever token the user happened to have, so the server on the other end can't tell which agent is calling. AI agent identity fixes that. Each agent gets its own short-lived, signed credential from an identity provider, presents it to every MCP server and API it calls, and the server checks the signature before it runs anything. In this guide you'll run a small local identity provider, an MCP server that verifies agent tokens against the provider's published keys, and a real agent whose whoami tool shows exactly which identity the server verified. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed. The cloud providers (Microsoft Entra ID, AWS IAM, Google Cloud, SPIFFE) are configured the same way, but they weren't run for this guide.
How do you give an AI agent its own identity?
Give it a workload identity, the same kind of non-human identity your services already use. An identity provider issues the agent a short-lived signed JWT, the agent presents it as a bearer token, and each server verifies it against the provider's public keys. With Promptise Foundry, the agent side is the identity argument of build_agent, plus an audience on each server:
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"support": HTTPServerSpec(
url="http://127.0.0.1:8310/mcp",
audience="api://support-tools",
),
},
identity=identity,The server side is JwksAuth:
server = MCPServer("support-tools", require_auth=True)
# Trust tokens from this issuer, minted for this server's audience, and nothing else.
server.add_middleware(
AuthMiddleware(
JwksAuth.from_discovery(
issuer="http://127.0.0.1:8311",
audience="api://support-tools",
)
)
)identity is an AgentIdentity backed by a credential provider, for example AgentIdentity.from_entra("support-bot", client_id=..., resource="api://support-tools") on Azure. When the agent connects, Promptise asks the provider for a token scoped to that server's audience and sends it as Authorization: Bearer …. The server fetches the issuer's public keys, checks the signature, issuer, audience and expiry, and hands your tools the verified identity on ctx.client. Authentication needs no shared secret on either side.
[02]
How agent identity works
Two identities meet in every request, and most mistakes come from mixing them up. The user is who asked. The agent is who acts.
Rendering diagram…
| User identity | Agent identity |
|---|---|---|
Answers | Who asked for this? | Which agent is acting? |
In Promptise | CallerContext, passed to each ainvoke call | AgentIdentity, passed once to build_agent |
Comes from | Your app's login | Your identity provider |
Used for | Conversation ownership, memory and cache isolation, approvals, observability | Authenticating to MCP servers and APIs, attribution, the server's audit log |
Reaches the MCP server in 1.2.1 | No | Yes, as the bearer token |
The last row is the one to remember. In this version, an MCP server sees the agent, not the user. If a tool must act as a specific person, the server can't learn who that is from the token; you'll see that tested near the end.
[03]
What you need
Python 3.10 or newer.
Promptise Foundry: pip install promptise. PyJWT and cryptography, which the server uses to check signatures, are installed with it.
An API key for a model provider. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.
No cloud account. A small local identity provider stands in for Entra, AWS or Google Cloud.
pip install promptise
export OPENAI_API_KEY="sk-..."[04]
Give your agent a verifiable identity, step by step
You'll build support-bot, a support agent that reads tickets from an MCP server. The server lets in only agents its identity provider vouches for, and only support-bot may read tickets. Three processes run on your machine: the identity provider on port 8311, the MCP server on 8310, and the agent.
Run a local identity provider
Someone has to vouch for the agent. On a cloud, that's Entra ID, AWS IAM or Google Cloud's metadata server. Locally, this script does the three jobs an OpenID Connect issuer does: it signs short-lived tokens, publishes a discovery document, and publishes its public keys.
"""A tiny OIDC issuer for local development.
It stands in for your cloud's identity service: it signs short-lived JWTs
for registered agents, and publishes its public keys so servers can check them.
Never run this in production.
"""
import json
import os
import time
import uuid
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import parse_qs, urlparse
import jwt
from cryptography.hazmat.primitives.asymmetric import rsa
ISSUER = "http://127.0.0.1:8311"
TOKEN_TTL = int(os.environ.get("TOKEN_TTL", "300")) # seconds
# The directory of agents this issuer vouches for: the system of record.
AGENTS = {
"support-bot": {"roles": ["support"]},
"reporting-bot": {"roles": ["reporting"]},
}
# A fresh signing key every start. Real issuers keep theirs in an HSM or KMS.
PRIVATE_KEY = rsa.generate_private_key(public_exponent=65537, key_size=2048)
KEY_ID = uuid.uuid4().hex[:8]
PUBLIC_JWK = {
**json.loads(jwt.algorithms.RSAAlgorithm.to_jwk(PRIVATE_KEY.public_key())),
"kid": KEY_ID,
"use": "sig",
"alg": "RS256",
}
def mint(agent: str, audience: str) -> str:
now = int(time.time())
claims = {
"iss": ISSUER,
"sub": agent,
"aud": audience,
"iat": now,
"exp": now + TOKEN_TTL,
"roles": AGENTS[agent]["roles"],
}
token = jwt.encode(claims, PRIVATE_KEY, algorithm="RS256", headers={"kid": KEY_ID})
print(f"[idp] minted token for {agent}, aud={audience}, expires in {TOKEN_TTL}s", flush=True)
return token
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
url = urlparse(self.path)
if url.path == "/.well-known/openid-configuration":
self.reply(200, {"issuer": ISSUER, "jwks_uri": f"{ISSUER}/jwks.json"})
elif url.path == "/jwks.json":
print("[idp] served public keys", flush=True)
self.reply(200, {"keys": [PUBLIC_JWK]})
elif url.path == "/token":
# A stand-in for a cloud metadata service. A real one knows which
# workload is asking; this one takes the agent's name as a parameter.
query = parse_qs(url.query)
agent = query.get("agent", [""])[0]
audience = query.get("audience", [""])[0]
if agent not in AGENTS or not audience:
self.reply(400, {"error": "unknown agent or missing audience"})
else:
self.reply(200, {"access_token": mint(agent, audience)})
else:
self.reply(404, {"error": "not found"})
def reply(self, status, body):
data = json.dumps(body).encode()
self.send_response(status)
self.send_header("content-type", "application/json")
self.send_header("content-length", str(len(data)))
self.end_headers()
self.wfile.write(data)
def log_message(self, *args):
pass # keep the output to the lines above
if __name__ == "__main__":
print(f"[idp] issuer {ISSUER}, key id {KEY_ID}, token lifetime {TOKEN_TTL}s", flush=True)
ThreadingHTTPServer(("127.0.0.1", 8311), Handler).serve_forever()Start it in its own terminal:
python local_idp.py[idp] issuer http://127.0.0.1:8311, key id f092d3b1, token lifetime 300sThree details carry over to the real thing. AGENTS is the directory, the one place where an agent exists, gets its roles and gets revoked. Every token has an aud (audience): the one server it's meant for. And each token expires after five minutes, so a leaked one is only useful briefly.
Verify agent tokens on the MCP server
A credential is only worth something if the server checks it. JwksAuth verifies each token against the issuer's public keys, so the server needs no secret of its own.
import os
import time
from promptise.mcp.server import (
AuditMiddleware,
AuthMiddleware,
JwksAuth,
MCPServer,
RequestContext,
RequireClientId,
)
server = MCPServer("support-tools", require_auth=True)
# Trust tokens from this issuer, minted for this server's audience, and nothing else.
server.add_middleware(
AuthMiddleware(
JwksAuth.from_discovery(
issuer="http://127.0.0.1:8311",
audience="api://support-tools",
)
)
)
server.add_middleware(AuditMiddleware("audit.jsonl", hmac_secret=os.environ["AUDIT_SECRET"]))
TICKETS = {
"T-101": {"subject": "Login fails after a password reset", "status": "open"},
}
@server.tool(read_only_hint=True)
async def whoami(ctx: RequestContext) -> dict:
"""Show which agent identity this server verified for the current call."""
client = ctx.client
return {
"client_id": client.client_id,
"issuer": client.issuer,
"audience": client.audience,
"roles": sorted(client.roles),
"token_expires_in_s": round(client.expires_at - time.time()),
}
@server.tool(read_only_hint=True, guards=[RequireClientId("support-bot")])
async def get_ticket(ticket_id: str) -> dict:
"""Get a support ticket by ID, for example "T-101"."""
ticket = TICKETS.get(ticket_id.strip().upper())
if ticket is None:
return {"error": f"No ticket {ticket_id}."}
return {"ticket_id": ticket_id.strip().upper(), **ticket}
if __name__ == "__main__":
server.run(transport="http", host="127.0.0.1", port=8310)Here's what each piece does:
`JwksAuth.from_discovery` reads the issuer's /.well-known/openid-configuration, refuses it if the document names a different issuer, and fetches the keys from the jwks_uri it lists. Every token must then carry a valid signature, iss equal to issuer, aud equal to audience, and an exp that hasn't passed.
`audience` is required. Without it, a token the same issuer minted for another server would work here too. That's called token substitution, and you'll try it in Step 5.
`require_auth=True` checks the token on every HTTP request, before the MCP layer sees it. A caller without a valid token gets 401 Unauthorized and never sees the tool list.
`ctx: RequestContext` gives a tool the verified caller on ctx.client: its client_id (the token's sub), issuer, audience, roles and expires_at, all read from a token whose signature checked out.
`RequireClientId("support-bot")` is authorization. Any agent the issuer vouches for can authenticate, but only support-bot may read tickets.
`AuditMiddleware` writes one signed line per tool call, including the verified identity.
Start the server in a second terminal:
AUDIT_SECRET=demo-audit-secret-not-for-prod python support_server.pyGive the agent an identity
An AgentIdentity names the agent, and a credential provider makes it verifiable. On a cloud you'd use a factory like AgentIdentity.from_entra. Here, CallableTokenProvider calls a function you write, which fetches a token from the local issuer:
import httpx
from promptise.identity import AgentIdentity, CallableTokenProvider
IDP = "http://127.0.0.1:8311"
DEFAULT_AUDIENCE = "api://support-tools"
def token_from_idp(audience: str | None) -> str:
"""Ask the issuer for a fresh JWT scoped to one audience."""
response = httpx.get(
f"{IDP}/token",
params={"agent": "support-bot", "audience": audience or DEFAULT_AUDIENCE},
)
response.raise_for_status()
return response.json()["access_token"]
identity = AgentIdentity(
"support-bot",
name="Support Bot",
owner="support-team",
credential=CallableTokenProvider(token_fn=token_from_idp, provider_label="local-idp"),
)Promptise calls token_fn with the audience a server asks for, or None when it wants the default. That's the same contract the cloud providers that mint on demand follow: Entra on a VM, AWS STS, Google Cloud's metadata server and the SPIFFE Workload API. If your issuer hands you one ready-made token instead, such as a CI job's OIDC token or a file Kubernetes keeps up to date, use AgentIdentity.from_oidc(issuer=..., token_file=...), or token_fn= or token_env_var= in place of token_file.
Look at the identity before you use it:
from support_identity import identity
print("repr: ", identity)
print("verifiable: ", identity.is_verifiable, "via", identity.credential_provider)
print("claims(): ", identity.claims())
print("idp_claims():", identity.idp_claims())
token = identity.get_credential("api://support-tools")
print("credential: ", token[:24] + "…", f"({len(token)} chars)")repr: AgentIdentity(agent_id='support-bot', name='Support Bot', owner='support-team', verifiable=True)
verifiable: True via local-idp
claims(): {'verifiable': True, 'agent_id': 'support-bot', 'name': 'Support Bot', 'owner': 'support-team', 'credential_provider': 'local-idp'}
idp_claims(): {'sub': 'support-bot', 'iss': 'http://127.0.0.1:8311', 'aud': 'api://support-tools'}
credential: eyJhbGciOiJSUzI1NiIsImtp… (581 chars)The printable parts never include the token: repr() and claims() carry identifiers only, so they're safe to log. idp_claims() reads who the issuer says the agent is, straight from the token. It doesn't check the signature, because the holder of a token isn't the one who needs convincing. The server is.
Connect the agent and ask who it is
Pass the identity to build_agent, and give the server spec the audience it expects:
import asyncio
from promptise import CallerContext, build_agent
from promptise.config import HTTPServerSpec
from support_identity import identity
async def main():
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"support": HTTPServerSpec(
url="http://127.0.0.1:8310/mcp",
audience="api://support-tools",
),
},
identity=identity,
instructions="You are a support assistant. Use your tools.",
trace_tools=True,
)
try:
print("Identity:", agent.identity)
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": "Which identity do the support tools see for you? Then get ticket T-101."}]},
caller=CallerContext(user_id="maya@shop.example"),
)
print(">>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())python agent.pyIdentity: AgentIdentity(agent_id='support-bot', name='Support Bot', owner='support-team', verifiable=True)
→ Invoking tool: whoami with {'payload': {}}
→ Invoking tool: get_ticket with {'ticket_id': 'T-101'}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 297}
✔ Tool result from get_ticket: {"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open"}
>>> I’m identified to the support tools as client_id "support-bot" (roles: ["support"]). I also retrieved ticket T-101:
…The server verified support-bot, issued by your issuer, for the api://support-tools audience, with the support role. Your agent code never handled a token: build_agent asked the identity for one scoped to the server's audience and sent it as the bearer token. The payload argument is how Promptise shows the model a tool that takes no parameters; the server ignores it.
The issuer's log shows the other half. One token was minted, and the server fetched the public keys once to check it:
[idp] issuer http://127.0.0.1:8311, key id f092d3b1, token lifetime 300s
[idp] minted token for support-bot, aud=api://support-tools, expires in 300s
[idp] served public keysAnd the server's audit log recorded who acted, from the verified token:
{"timestamp": 1791654578.314502, "tool": "whoami", "client_id": "support-bot", "request_id": "469dc0a206af", "status": "ok", "duration_s": 0.0, "identity": {"subject": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"]}, …Notice who's missing. The agent was working for maya@shop.example, and neither the tool result nor the audit line mentions her. The server knows which agent called, not which user asked.
Try to get in as someone else
A check you haven't seen fail is a check you're trusting on faith. This script knocks on the server with every kind of wrong credential:
"""Knock on the support server with every kind of wrong credential."""
import asyncio
import time
import httpx
import jwt
from cryptography.hazmat.primitives.asymmetric import rsa
from promptise.mcp.client import MCPClient
SERVER = "http://127.0.0.1:8310/mcp"
IDP = "http://127.0.0.1:8311"
def real_token(agent: str, audience: str) -> str:
response = httpx.get(f"{IDP}/token", params={"agent": agent, "audience": audience})
return response.json()["access_token"]
def forged_token() -> str:
"""Claims to be support-bot, but is signed with the attacker's own key."""
kid = httpx.get(f"{IDP}/jwks.json").json()["keys"][0]["kid"] # the key id is public
attacker_key = rsa.generate_private_key(public_exponent=65537, key_size=2048)
now = int(time.time())
claims = {"iss": IDP, "sub": "support-bot", "aud": "api://support-tools", "iat": now, "exp": now + 300}
return jwt.encode(claims, attacker_key, algorithm="RS256", headers={"kid": kid})
async def attempt(label: str, token: str | None) -> None:
try:
async with MCPClient(url=SERVER, bearer_token=token) as client:
result = await client.call_tool("get_ticket", {"ticket_id": "T-101"})
text = " ".join(result.content[0].text.split())
print(f"{label}: connected, get_ticket -> {text[:150]}")
except Exception as exc:
print(f"{label}: {type(exc).__name__}: {exc}")
async def main():
await attempt("no token", None)
await attempt("forged signature", forged_token())
await attempt("token for another server", real_token("support-bot", "api://billing"))
await attempt("another agent", real_token("reporting-bot", "api://support-tools"))
await attempt("support-bot", real_token("support-bot", "api://support-tools"))
asyncio.run(main())no token: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
forged signature: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
token for another server: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
another agent: connected, get_ticket -> { "error": { "code": "ACCESS_DENIED", "message": "Client 'reporting-bot' is not in the allowed list [support-bot]", "retryable": false, "details": { "
support-bot: connected, get_ticket -> {"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open"}Each line is one layer doing its job:
No token never reaches MCP at all.
A forged token names the right agent, audience and key id, but its signature doesn't match the issuer's published key.
A real token for another server is the subtle one. It's genuinely support-bot, signed by the right issuer, but minted for api://billing. Without the audience check, any server that trusts this issuer could replay it here.
Another agent has a perfectly valid identity, so it authenticates. RequireClientId then refuses it the tool. Authentication says who you are; guards say what you may do.
[05]
One identity, several servers
Real agents call more than one service, and each should get its own token. That way a token stolen from one server can't be used against another. Set audience on each HTTPServerSpec, and one identity serves them all. Here's a second server that only accepts api://billing:
server = MCPServer("billing", require_auth=True)
server.add_middleware(
AuthMiddleware(JwksAuth.from_discovery(issuer="http://127.0.0.1:8311", audience="api://billing"))
)Its tool is called billing_whoami, so its name doesn't clash with the support server's whoami. Start it on port 8312, then connect the agent to both:
agent = await build_agent(
model="openai:gpt-5-mini",
servers={
"support": HTTPServerSpec(url="http://127.0.0.1:8310/mcp", audience="api://support-tools"),
"billing": HTTPServerSpec(url="http://127.0.0.1:8312/mcp", audience="api://billing"),
},
identity=identity,
trace_tools=True,
)→ Invoking tool: whoami with {'payload': {}}
→ Invoking tool: billing_whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 298}
✔ Tool result from billing_whoami: {"client_id": "support-bot", "audience": "api://billing", "token_expires_in_s": 298}
>>> I called both services. Results (relevant fields):
…The issuer minted one token per audience:
[idp] minted token for support-bot, aud=api://support-tools, expires in 300s
[idp] minted token for support-bot, aud=api://billing, expires in 300s
[idp] served public keys[06]
How the credential is cached and refreshed
Asking the issuer for a new token on every call would be slow and noisy. The identity's provider caches one token per audience, reads the expiry from the token's exp claim, and fetches a new one when the cached token is within 60 seconds of expiring. To watch it happen in under a minute, restart the issuer with 75-second tokens:
TOKEN_TTL=75 python local_idp.py"""Watch the credential cache: reuse, refresh before expiry, one token per audience."""
import time
from promptise.identity import decode_jwt_expiry
from support_identity import identity
start = time.time()
previous = {}
for at, audience in [(0, "api://support-tools"), (1, "api://billing"), (5, "api://support-tools"), (20, "api://support-tools")]:
time.sleep(max(0, start + at - time.time()))
token = identity.get_credential(audience)
status = "reused" if previous.get(audience) == token else "new token"
previous[audience] = token
left = decode_jwt_expiry(token) - time.time()
print(f"t={at:>2}s {audience:<20} {status:<9} expires in {left:.0f}s")t= 0s api://support-tools new token expires in 75s
t= 1s api://billing new token expires in 74s
t= 5s api://support-tools reused expires in 70s
t=20s api://support-tools new token expires in 74sAt five seconds the cached token still had 70 seconds left, so it was reused. At twenty seconds it had 55 left, inside the 60-second window, so the identity fetched a fresh one. The billing token lives in its own cache entry. The cache is in memory and per process: a token is never written to disk, and two processes each fetch their own.
[07]
Swap in Entra ID, AWS, Google Cloud or SPIFFE
The local issuer was a stand-in. In production, the platform your agent runs on is the issuer, and the agent picks up its identity from the environment: a metadata service, or a token file the platform keeps fresh. You don't store a secret anywhere. Only the line that builds the identity changes; the server and build_agent stay exactly as they are.
"""The same agent, backed by each cloud provider. Constructed here, not run against a cloud."""
from promptise.identity import AgentIdentity
identities = {
"Microsoft Entra ID": AgentIdentity.from_entra(
"support-bot", client_id="<managed-identity-client-id>", resource="api://support-tools"
),
"AWS IAM": AgentIdentity.from_aws("support-bot", region="eu-central-1", audience="api://support-tools"),
"Google Cloud": AgentIdentity.from_gcp("support-bot", audience="api://support-tools"),
"SPIFFE / SPIRE": AgentIdentity.from_spiffe(
"support-bot", token_file="/run/spiffe/jwt_svid.token", audience="api://support-tools"
),
"Generic OIDC": AgentIdentity.from_oidc(
"support-bot", issuer="https://token.actions.githubusercontent.com", token_env_var="AGENT_OIDC_TOKEN"
),
}
for platform, identity in identities.items():
print(f"{platform:<20} {identity.credential_provider}")Microsoft Entra ID entra-imds
AWS IAM aws-sts
Google Cloud gcp-metadata
SPIFFE / SPIRE spiffe-file
Generic OIDC oidc:https://token.actions.githubusercontent.comThis confirms each factory accepts these settings and picks the mode you'd expect off-platform. It doesn't prove the live round trip: none of these were run against a real cloud for this guide. Promptise's Security page is just as frank, and points to opt-in integration tests you can run inside each platform before you depend on one. On the server side, point JwksAuth.from_discovery at your platform's issuer URL instead of the local one.
Where the agent runs | Factory | Where the token comes from |
|---|---|---|
Azure VM, Container Apps | from_entra() | Managed identity, through the instance metadata service |
Azure Kubernetes Service | from_entra() | The file in AZURE_FEDERATED_TOKEN_FILE |
AWS Lambda, EC2, ECS | from_aws() | STS GetWebIdentityToken, which needs boto3 |
AWS EKS | from_aws(mode="projected") | A projected token file, read from PROMPTISE_IDENTITY_TOKEN_FILE |
Google Compute Engine, Cloud Run, GKE | from_gcp() | The metadata server |
A SPIFFE or SPIRE mesh | from_spiffe() | The Workload API socket, or a spiffe-helper file |
GitHub Actions, GitLab CI, any OIDC issuer | from_oidc() | A file, an environment variable or your own function |
AgentIdentity.auto("support-bot") picks the provider for you from environment variables, and raises PlatformDetectionError when it finds none, for example on your laptop. Each provider has its own setup page under Credential providers.
[08]
Long-running agents and token expiry
build_agent fetches each server's token once, when it connects. The identity keeps refreshing its own cache, but the open connection keeps sending the token it started with. An agent that lives longer than its token will eventually be turned away. This run used 75-second tokens and waited past their expiry, plus the server's 60-second grace period:
start = time.time()
await agent.ainvoke(QUESTION)
await asyncio.sleep(140) # longer than the token lifetime plus the server's 60s leeway
print(f"--- {time.time() - start:.0f}s later")
fresh = identity.get_credential("api://support-tools")
print(f"The identity's own cached token now expires in {decode_jwt_expiry(fresh) - time.time():.0f}s")
try:
await asyncio.wait_for(agent.ainvoke(QUESTION), timeout=120)
except TimeoutError:
print("No answer after 120s: the tool call is still waiting.")→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 72}
--- 142s later
The identity's own cached token now expires in 75s
→ Invoking tool: whoami with {'payload': {}}
Server at http://127.0.0.1:8310/mcp answered HTTP 401 Unauthorized
No answer after 120s: the tool call is still waiting.Two problems show here. The identity had a fresh token ready, but the connection never used it, so the server rejected the call. Worse, the rejected call didn't fail: it waited until the timeout gave up on it. Without asyncio.wait_for, an earlier run waited more than seven minutes before it was stopped by hand.
Until that changes, build a short-lived agent per job or per session, from the same identity. Each build connects with a token that has at least a minute left, and the shared identity means its cache is shared too:
async def run_job(question: str) -> str:
agent = await build_agent(
model="openai:gpt-5-mini",
servers={"support": HTTPServerSpec(url="http://127.0.0.1:8310/mcp", audience="api://support-tools")},
identity=identity, # the same identity, so its credential cache is shared
trace_tools=True,
)
try:
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
return result["messages"][-1].content
finally:
await agent.shutdown()→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 72}
>>> The whoami response shows token_expires_in_s = 72 (seconds).
--- 144s later
→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 73}
>>> I called whoami. The returned token_expires_in_s value is 73 — the token will expire in 73 seconds from the time of the call.The second job got a new token and went through. Also put a timeout around ainvoke in anything long-running, and pick a token lifetime comfortably longer than your longest job.
[09]
The user's token stays in the agent
CallerContext has a bearer_token field, and it's tempting to expect the user's token to travel to the MCP server with each call. In 1.2.1 it doesn't; the field's own docstring calls MCP token forwarding "a planned enhancement". This check gives the agent no identity, passes a valid token as the caller's, and calls a server that checks tokens per tool:
caller = CallerContext(user_id="maya@shop.example", bearer_token=token_from_idp("api://support-tools"))
await agent.ainvoke({"messages": [{"role": "user", "content": "Call whoami once."}]}, caller=caller)→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {
"error": {
"code": "AUTHENTICATION_ERROR",
"message": "Missing authentication token",
"retryable": false
}
}So authenticate the agent with AgentIdentity, and keep user-level decisions where the user is known: in your app, in the agent's CallerContext for isolation and approvals, or as an explicit tool argument the server validates against its own records.
[10]
Honest limits
These are true of Promptise Foundry 1.2.1:
The token is fixed per connection, and a call made after it expires waits instead of failing, as shown above. Build an agent per job and wrap ainvoke in a timeout.
The user's token isn't forwarded to MCP servers, as shown above. A server sees the agent's identity only.
If the issuer is down when the agent connects, build_agent logs Agent identity 'support-bot' could not acquire a credential for MCP server audience 'api://support-tools' (…); connecting without it. and the server then refuses the connection with a 401. That's the safe direction, but watch for the warning in your logs.
An explicit `bearer_token` on an `HTTPServerSpec` wins over the identity for that server, so remove it when you switch to agent identity.
The server allows 60 seconds of clock skew. A token is accepted for up to a minute after its exp. JwksAuth(jwks_url=..., audience=..., leeway=...) lets you change that; JwksAuth.from_discovery doesn't take a leeway argument.
A new signing key is picked up automatically, but not instantly. When a token names a key id the server hasn't seen, it fetches the issuer's keys again, but only if its last fetch was more than 30 seconds ago. In this guide's runs, a token signed with a brand-new key was refused inside that window and accepted after it. Most issuers publish a new key before they start signing with it, which avoids the gap.
`subject()` and `idp_claims()` don't verify the token. They're for attribution inside your own process. Proof comes from the server checking the signature.
The cloud providers weren't run here. Run Promptise's opt-in integration tests on your platform before you rely on one.
[11]
Frequently asked questions
What is AI agent identity?
It's an identity for the agent itself, separate from the person using it: a name in your identity provider, plus a short-lived signed credential the agent presents to every service it calls. Services verify the credential, authorize the agent by name or role, and record which agent did what. In Promptise Foundry it's AgentIdentity, passed to build_agent(identity=...).
How do you authenticate an AI agent to an MCP server?
Have the agent present a signed JWT from your identity provider as a bearer token, and verify it on the server against the provider's public keys. In Promptise, give the agent an AgentIdentity and set audience on each HTTPServerSpec; on the server, add AuthMiddleware(JwksAuth.from_discovery(issuer=..., audience=...)) and require_auth=True. Then use guards like RequireClientId or HasRole to decide what each agent may do.
What is a non-human identity for AI agents?
A non-human identity is an identity for software rather than a person, like a service account, a managed identity or an IAM role. AI agents are a new kind of non-human identity: they act on their own, often all day, so they need credentials that are short-lived, scoped to one resource, and revocable in one place. Workload identity platforms such as Entra ID, AWS IAM, Google Cloud and SPIFFE already issue exactly that.
Do I still need user authentication if the agent has an identity?
Yes. The agent's identity proves which agent is calling; it says nothing about which person asked. Keep authenticating users in your app, pass who they are to the agent with CallerContext, and decide user-level permissions there, because in 1.2.1 the MCP server only sees the agent.
Does the agent need an API key or client secret?
Not on a cloud platform. The agent gets its token from the platform's metadata service or a token file the platform rotates, so there's no long-lived secret to store or leak. The server needs none either, because it verifies tokens with the issuer's public keys.
[12]
Where to go next
Agent Identity overview: local and verifiable identities, and where the identity shows up.
Agent Identity guide: an end-to-end scenario with delegation between agents.
Architecture and Security: credential lifecycle, per-resource credentials and the threat model.
Generic OIDC provider: from_oidc with files, functions and environment variables.
Authentication & Security: JwksAuth, JWTAuth, guards and ClientContext on the server.
Server configuration: every option on HTTPServerSpec, including audience.
How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.
OpenAPI to MCP: Turn Any REST API into an MCP Server: put an existing API behind an MCP server, then protect it the same way.