PromptisePromptise
Docs
GitHub
Promptise - AI Framework LogoPromptise

The foundation layer for agentic intelligence. Build, secure and operate autonomous AI systems with Promptise Foundry.

pip install promptise

[01] Foundry

  • MCPcast
  • The Promptise Agent
  • Reasoning Engine
  • MCP
  • Agent Runtime
  • Prompt Engineering
  • Execution Engine
  • Agent Identity

[02] Resources

  • Documentation
  • GitHub
  • Guides
  • Learning Paths
  • Questions

[03] Company

  • About
  • Terms of Service
  • Privacy Policy
  • Cookie Policy
  • Subprocessors

© 2026 Promptise by Manser Ventures. All rights reserved.

Open source · Python

← Guides> AI Engineering

AI Agent Identity: Authenticate Agents to MCP Servers

Give your AI agent its own verifiable identity: short-lived JWTs from your IdP, checked by MCP servers with JWKS. Real runs, refresh and honest limits.

Level
Advanced
Reading time
22 min
Published
Oct 10, 2026
By
Promptise Team
  • AI Agent Identity
  • Authentication
  • MCP
  • Workload Identity
  • Python
  • Promptise Foundry

Most AI agents call their tools with a shared API key, or with whatever token the user happened to have, so the server on the other end can't tell which agent is calling. AI agent identity fixes that. Each agent gets its own short-lived, signed credential from an identity provider, presents it to every MCP server and API it calls, and the server checks the signature before it runs anything. In this guide you'll run a small local identity provider, an MCP server that verifies agent tokens against the provider's published keys, and a real agent whose whoami tool shows exactly which identity the server verified. Every snippet ran against Promptise Foundry 1.2.1, and every output is what it printed. The cloud providers (Microsoft Entra ID, AWS IAM, Google Cloud, SPIFFE) are configured the same way, but they weren't run for this guide.

[01]

How do you give an AI agent its own identity?

Give it a workload identity, the same kind of non-human identity your services already use. An identity provider issues the agent a short-lived signed JWT, the agent presents it as a bearer token, and each server verifies it against the provider's public keys. With Promptise Foundry, the agent side is the identity argument of build_agent, plus an audience on each server:

Pythonagent.py
1
2
3
4
5
6
7
8
9
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "support": HTTPServerSpec(
                url="http://127.0.0.1:8310/mcp",
                audience="api://support-tools",
            ),
        },
        identity=identity,

The server side is JwksAuth:

Pythonsupport_server.py
1
2
3
4
5
6
7
8
9
10
11
server = MCPServer("support-tools", require_auth=True)

# Trust tokens from this issuer, minted for this server's audience, and nothing else.
server.add_middleware(
    AuthMiddleware(
        JwksAuth.from_discovery(
            issuer="http://127.0.0.1:8311",
            audience="api://support-tools",
        )
    )
)

identity is an AgentIdentity backed by a credential provider, for example AgentIdentity.from_entra("support-bot", client_id=..., resource="api://support-tools") on Azure. When the agent connects, Promptise asks the provider for a token scoped to that server's audience and sends it as Authorization: Bearer …. The server fetches the issuer's public keys, checks the signature, issuer, audience and expiry, and hands your tools the verified identity on ctx.client. Authentication needs no shared secret on either side.


[02]

How agent identity works

Two identities meet in every request, and most mistakes come from mixing them up. The user is who asked. The agent is who acts.

Rendering diagram…

User identity

Agent identity

Answers

Who asked for this?

Which agent is acting?

In Promptise

CallerContext, passed to each ainvoke call

AgentIdentity, passed once to build_agent

Comes from

Your app's login

Your identity provider

Used for

Conversation ownership, memory and cache isolation, approvals, observability

Authenticating to MCP servers and APIs, attribution, the server's audit log

Reaches the MCP server in 1.2.1

No

Yes, as the bearer token

The last row is the one to remember. In this version, an MCP server sees the agent, not the user. If a tool must act as a specific person, the server can't learn who that is from the token; you'll see that tested near the end.


[03]

What you need

  • Python 3.10 or newer.

  • Promptise Foundry: pip install promptise. PyJWT and cryptography, which the server uses to check signatures, are installed with it.

  • An API key for a model provider. This guide uses OpenAI's gpt-5-mini; other providers work by changing the model string, as listed in Models & Providers.

  • No cloud account. A small local identity provider stands in for Entra, AWS or Google Cloud.

>_Terminal
pip install promptise
export OPENAI_API_KEY="sk-..."

[04]

Give your agent a verifiable identity, step by step

You'll build support-bot, a support agent that reads tickets from an MCP server. The server lets in only agents its identity provider vouches for, and only support-bot may read tickets. Three processes run on your machine: the identity provider on port 8311, the MCP server on 8310, and the agent.

Step 01

Run a local identity provider

Someone has to vouch for the agent. On a cloud, that's Entra ID, AWS IAM or Google Cloud's metadata server. Locally, this script does the three jobs an OpenID Connect issuer does: it signs short-lived tokens, publishes a discovery document, and publishes its public keys.

Pythonlocal_idp.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
"""A tiny OIDC issuer for local development.

It stands in for your cloud's identity service: it signs short-lived JWTs
for registered agents, and publishes its public keys so servers can check them.
Never run this in production.
"""

import json
import os
import time
import uuid
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import parse_qs, urlparse

import jwt
from cryptography.hazmat.primitives.asymmetric import rsa

ISSUER = "http://127.0.0.1:8311"
TOKEN_TTL = int(os.environ.get("TOKEN_TTL", "300"))  # seconds

# The directory of agents this issuer vouches for: the system of record.
AGENTS = {
    "support-bot": {"roles": ["support"]},
    "reporting-bot": {"roles": ["reporting"]},
}

# A fresh signing key every start. Real issuers keep theirs in an HSM or KMS.
PRIVATE_KEY = rsa.generate_private_key(public_exponent=65537, key_size=2048)
KEY_ID = uuid.uuid4().hex[:8]
PUBLIC_JWK = {
    **json.loads(jwt.algorithms.RSAAlgorithm.to_jwk(PRIVATE_KEY.public_key())),
    "kid": KEY_ID,
    "use": "sig",
    "alg": "RS256",
}


def mint(agent: str, audience: str) -> str:
    now = int(time.time())
    claims = {
        "iss": ISSUER,
        "sub": agent,
        "aud": audience,
        "iat": now,
        "exp": now + TOKEN_TTL,
        "roles": AGENTS[agent]["roles"],
    }
    token = jwt.encode(claims, PRIVATE_KEY, algorithm="RS256", headers={"kid": KEY_ID})
    print(f"[idp] minted token for {agent}, aud={audience}, expires in {TOKEN_TTL}s", flush=True)
    return token


class Handler(BaseHTTPRequestHandler):
    def do_GET(self):
        url = urlparse(self.path)
        if url.path == "/.well-known/openid-configuration":
            self.reply(200, {"issuer": ISSUER, "jwks_uri": f"{ISSUER}/jwks.json"})
        elif url.path == "/jwks.json":
            print("[idp] served public keys", flush=True)
            self.reply(200, {"keys": [PUBLIC_JWK]})
        elif url.path == "/token":
            # A stand-in for a cloud metadata service. A real one knows which
            # workload is asking; this one takes the agent's name as a parameter.
            query = parse_qs(url.query)
            agent = query.get("agent", [""])[0]
            audience = query.get("audience", [""])[0]
            if agent not in AGENTS or not audience:
                self.reply(400, {"error": "unknown agent or missing audience"})
            else:
                self.reply(200, {"access_token": mint(agent, audience)})
        else:
            self.reply(404, {"error": "not found"})

    def reply(self, status, body):
        data = json.dumps(body).encode()
        self.send_response(status)
        self.send_header("content-type", "application/json")
        self.send_header("content-length", str(len(data)))
        self.end_headers()
        self.wfile.write(data)

    def log_message(self, *args):
        pass  # keep the output to the lines above


if __name__ == "__main__":
    print(f"[idp] issuer {ISSUER}, key id {KEY_ID}, token lifetime {TOKEN_TTL}s", flush=True)
    ThreadingHTTPServer(("127.0.0.1", 8311), Handler).serve_forever()

Start it in its own terminal:

>_Terminal
python local_idp.py
Output
[idp] issuer http://127.0.0.1:8311, key id f092d3b1, token lifetime 300s

Three details carry over to the real thing. AGENTS is the directory, the one place where an agent exists, gets its roles and gets revoked. Every token has an aud (audience): the one server it's meant for. And each token expires after five minutes, so a leaked one is only useful briefly.

Step 02

Verify agent tokens on the MCP server

A credential is only worth something if the server checks it. JwksAuth verifies each token against the issuer's public keys, so the server needs no secret of its own.

Pythonsupport_server.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
import os
import time

from promptise.mcp.server import (
    AuditMiddleware,
    AuthMiddleware,
    JwksAuth,
    MCPServer,
    RequestContext,
    RequireClientId,
)

server = MCPServer("support-tools", require_auth=True)

# Trust tokens from this issuer, minted for this server's audience, and nothing else.
server.add_middleware(
    AuthMiddleware(
        JwksAuth.from_discovery(
            issuer="http://127.0.0.1:8311",
            audience="api://support-tools",
        )
    )
)
server.add_middleware(AuditMiddleware("audit.jsonl", hmac_secret=os.environ["AUDIT_SECRET"]))

TICKETS = {
    "T-101": {"subject": "Login fails after a password reset", "status": "open"},
}


@server.tool(read_only_hint=True)
async def whoami(ctx: RequestContext) -> dict:
    """Show which agent identity this server verified for the current call."""
    client = ctx.client
    return {
        "client_id": client.client_id,
        "issuer": client.issuer,
        "audience": client.audience,
        "roles": sorted(client.roles),
        "token_expires_in_s": round(client.expires_at - time.time()),
    }


@server.tool(read_only_hint=True, guards=[RequireClientId("support-bot")])
async def get_ticket(ticket_id: str) -> dict:
    """Get a support ticket by ID, for example "T-101"."""
    ticket = TICKETS.get(ticket_id.strip().upper())
    if ticket is None:
        return {"error": f"No ticket {ticket_id}."}
    return {"ticket_id": ticket_id.strip().upper(), **ticket}


if __name__ == "__main__":
    server.run(transport="http", host="127.0.0.1", port=8310)

Here's what each piece does:

  • `JwksAuth.from_discovery` reads the issuer's /.well-known/openid-configuration, refuses it if the document names a different issuer, and fetches the keys from the jwks_uri it lists. Every token must then carry a valid signature, iss equal to issuer, aud equal to audience, and an exp that hasn't passed.

  • `audience` is required. Without it, a token the same issuer minted for another server would work here too. That's called token substitution, and you'll try it in Step 5.

  • `require_auth=True` checks the token on every HTTP request, before the MCP layer sees it. A caller without a valid token gets 401 Unauthorized and never sees the tool list.

  • `ctx: RequestContext` gives a tool the verified caller on ctx.client: its client_id (the token's sub), issuer, audience, roles and expires_at, all read from a token whose signature checked out.

  • `RequireClientId("support-bot")` is authorization. Any agent the issuer vouches for can authenticate, but only support-bot may read tickets.

  • `AuditMiddleware` writes one signed line per tool call, including the verified identity.

Start the server in a second terminal:

>_Terminal
AUDIT_SECRET=demo-audit-secret-not-for-prod python support_server.py
Step 03

Give the agent an identity

An AgentIdentity names the agent, and a credential provider makes it verifiable. On a cloud you'd use a factory like AgentIdentity.from_entra. Here, CallableTokenProvider calls a function you write, which fetches a token from the local issuer:

Pythonsupport_identity.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import httpx

from promptise.identity import AgentIdentity, CallableTokenProvider

IDP = "http://127.0.0.1:8311"
DEFAULT_AUDIENCE = "api://support-tools"


def token_from_idp(audience: str | None) -> str:
    """Ask the issuer for a fresh JWT scoped to one audience."""
    response = httpx.get(
        f"{IDP}/token",
        params={"agent": "support-bot", "audience": audience or DEFAULT_AUDIENCE},
    )
    response.raise_for_status()
    return response.json()["access_token"]


identity = AgentIdentity(
    "support-bot",
    name="Support Bot",
    owner="support-team",
    credential=CallableTokenProvider(token_fn=token_from_idp, provider_label="local-idp"),
)

Promptise calls token_fn with the audience a server asks for, or None when it wants the default. That's the same contract the cloud providers that mint on demand follow: Entra on a VM, AWS STS, Google Cloud's metadata server and the SPIFFE Workload API. If your issuer hands you one ready-made token instead, such as a CI job's OIDC token or a file Kubernetes keeps up to date, use AgentIdentity.from_oidc(issuer=..., token_file=...), or token_fn= or token_env_var= in place of token_file.

Look at the identity before you use it:

Pythoninspect_identity.py
1
2
3
4
5
6
7
8
from support_identity import identity

print("repr:       ", identity)
print("verifiable: ", identity.is_verifiable, "via", identity.credential_provider)
print("claims():   ", identity.claims())
print("idp_claims():", identity.idp_claims())
token = identity.get_credential("api://support-tools")
print("credential: ", token[:24] + "…", f"({len(token)} chars)")
Output
repr:        AgentIdentity(agent_id='support-bot', name='Support Bot', owner='support-team', verifiable=True)
verifiable:  True via local-idp
claims():    {'verifiable': True, 'agent_id': 'support-bot', 'name': 'Support Bot', 'owner': 'support-team', 'credential_provider': 'local-idp'}
idp_claims(): {'sub': 'support-bot', 'iss': 'http://127.0.0.1:8311', 'aud': 'api://support-tools'}
credential:  eyJhbGciOiJSUzI1NiIsImtp… (581 chars)

The printable parts never include the token: repr() and claims() carry identifiers only, so they're safe to log. idp_claims() reads who the issuer says the agent is, straight from the token. It doesn't check the signature, because the holder of a token isn't the one who needs convincing. The server is.

Step 04

Connect the agent and ask who it is

Pass the identity to build_agent, and give the server spec the audience it expects:

Pythonagent.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
import asyncio

from promptise import CallerContext, build_agent
from promptise.config import HTTPServerSpec

from support_identity import identity


async def main():
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "support": HTTPServerSpec(
                url="http://127.0.0.1:8310/mcp",
                audience="api://support-tools",
            ),
        },
        identity=identity,
        instructions="You are a support assistant. Use your tools.",
        trace_tools=True,
    )
    try:
        print("Identity:", agent.identity)
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Which identity do the support tools see for you? Then get ticket T-101."}]},
            caller=CallerContext(user_id="maya@shop.example"),
        )
        print(">>>", result["messages"][-1].content)
    finally:
        await agent.shutdown()


asyncio.run(main())
>_Terminal
python agent.py
Output
Identity: AgentIdentity(agent_id='support-bot', name='Support Bot', owner='support-team', verifiable=True)
→ Invoking tool: whoami with {'payload': {}}
→ Invoking tool: get_ticket with {'ticket_id': 'T-101'}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 297}
✔ Tool result from get_ticket: {"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open"}
>>> I’m identified to the support tools as client_id "support-bot" (roles: ["support"]). I also retrieved ticket T-101:
…

The server verified support-bot, issued by your issuer, for the api://support-tools audience, with the support role. Your agent code never handled a token: build_agent asked the identity for one scoped to the server's audience and sent it as the bearer token. The payload argument is how Promptise shows the model a tool that takes no parameters; the server ignores it.

The issuer's log shows the other half. One token was minted, and the server fetched the public keys once to check it:

Output
[idp] issuer http://127.0.0.1:8311, key id f092d3b1, token lifetime 300s
[idp] minted token for support-bot, aud=api://support-tools, expires in 300s
[idp] served public keys

And the server's audit log recorded who acted, from the verified token:

Output
{"timestamp": 1791654578.314502, "tool": "whoami", "client_id": "support-bot", "request_id": "469dc0a206af", "status": "ok", "duration_s": 0.0, "identity": {"subject": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"]}, …

Notice who's missing. The agent was working for maya@shop.example, and neither the tool result nor the audit line mentions her. The server knows which agent called, not which user asked.

Step 05

Try to get in as someone else

A check you haven't seen fail is a check you're trusting on faith. This script knocks on the server with every kind of wrong credential:

Pythonimpostors.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
"""Knock on the support server with every kind of wrong credential."""

import asyncio
import time

import httpx
import jwt
from cryptography.hazmat.primitives.asymmetric import rsa

from promptise.mcp.client import MCPClient

SERVER = "http://127.0.0.1:8310/mcp"
IDP = "http://127.0.0.1:8311"


def real_token(agent: str, audience: str) -> str:
    response = httpx.get(f"{IDP}/token", params={"agent": agent, "audience": audience})
    return response.json()["access_token"]


def forged_token() -> str:
    """Claims to be support-bot, but is signed with the attacker's own key."""
    kid = httpx.get(f"{IDP}/jwks.json").json()["keys"][0]["kid"]  # the key id is public
    attacker_key = rsa.generate_private_key(public_exponent=65537, key_size=2048)
    now = int(time.time())
    claims = {"iss": IDP, "sub": "support-bot", "aud": "api://support-tools", "iat": now, "exp": now + 300}
    return jwt.encode(claims, attacker_key, algorithm="RS256", headers={"kid": kid})


async def attempt(label: str, token: str | None) -> None:
    try:
        async with MCPClient(url=SERVER, bearer_token=token) as client:
            result = await client.call_tool("get_ticket", {"ticket_id": "T-101"})
            text = " ".join(result.content[0].text.split())
            print(f"{label}: connected, get_ticket -> {text[:150]}")
    except Exception as exc:
        print(f"{label}: {type(exc).__name__}: {exc}")


async def main():
    await attempt("no token", None)
    await attempt("forged signature", forged_token())
    await attempt("token for another server", real_token("support-bot", "api://billing"))
    await attempt("another agent", real_token("reporting-bot", "api://support-tools"))
    await attempt("support-bot", real_token("support-bot", "api://support-tools"))


asyncio.run(main())
Output
no token: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
forged signature: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
token for another server: MCPConnectionRejectedError: Server at http://127.0.0.1:8310/mcp rejected the connection: 401 Unauthorized. Check the bearer_token/api_key configured for it.
another agent: connected, get_ticket -> { "error": { "code": "ACCESS_DENIED", "message": "Client 'reporting-bot' is not in the allowed list [support-bot]", "retryable": false, "details": { "
support-bot: connected, get_ticket -> {"ticket_id": "T-101", "subject": "Login fails after a password reset", "status": "open"}

Each line is one layer doing its job:

  • No token never reaches MCP at all.

  • A forged token names the right agent, audience and key id, but its signature doesn't match the issuer's published key.

  • A real token for another server is the subtle one. It's genuinely support-bot, signed by the right issuer, but minted for api://billing. Without the audience check, any server that trusts this issuer could replay it here.

  • Another agent has a perfectly valid identity, so it authenticates. RequireClientId then refuses it the tool. Authentication says who you are; guards say what you may do.


[05]

One identity, several servers

Real agents call more than one service, and each should get its own token. That way a token stolen from one server can't be used against another. Set audience on each HTTPServerSpec, and one identity serves them all. Here's a second server that only accepts api://billing:

Pythonbilling_server.py
server = MCPServer("billing", require_auth=True)
server.add_middleware(
    AuthMiddleware(JwksAuth.from_discovery(issuer="http://127.0.0.1:8311", audience="api://billing"))
)

Its tool is called billing_whoami, so its name doesn't clash with the support server's whoami. Start it on port 8312, then connect the agent to both:

Pythontwo_servers.py
1
2
3
4
5
6
7
8
9
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={
            "support": HTTPServerSpec(url="http://127.0.0.1:8310/mcp", audience="api://support-tools"),
            "billing": HTTPServerSpec(url="http://127.0.0.1:8312/mcp", audience="api://billing"),
        },
        identity=identity,
        trace_tools=True,
    )
Output
→ Invoking tool: whoami with {'payload': {}}
→ Invoking tool: billing_whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 298}
✔ Tool result from billing_whoami: {"client_id": "support-bot", "audience": "api://billing", "token_expires_in_s": 298}
>>> I called both services. Results (relevant fields):
…

The issuer minted one token per audience:

Output
[idp] minted token for support-bot, aud=api://support-tools, expires in 300s
[idp] minted token for support-bot, aud=api://billing, expires in 300s
[idp] served public keys
Note

Per-audience tokens need a provider that can mint on demand: Entra on a VM, AWS STS, Google Cloud's metadata server, the SPIFFE Workload API, or your own CallableTokenProvider. Token files and from_oidc hold one token with a fixed audience and ignore audience. For a second audience from such a source, issue a second token at the source and build a second identity.


[06]

How the credential is cached and refreshed

Asking the issuer for a new token on every call would be slow and noisy. The identity's provider caches one token per audience, reads the expiry from the token's exp claim, and fetches a new one when the cached token is within 60 seconds of expiring. To watch it happen in under a minute, restart the issuer with 75-second tokens:

>_Terminal
TOKEN_TTL=75 python local_idp.py
Pythonrefresh_demo.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
"""Watch the credential cache: reuse, refresh before expiry, one token per audience."""

import time

from promptise.identity import decode_jwt_expiry

from support_identity import identity

start = time.time()
previous = {}
for at, audience in [(0, "api://support-tools"), (1, "api://billing"), (5, "api://support-tools"), (20, "api://support-tools")]:
    time.sleep(max(0, start + at - time.time()))
    token = identity.get_credential(audience)
    status = "reused" if previous.get(audience) == token else "new token"
    previous[audience] = token
    left = decode_jwt_expiry(token) - time.time()
    print(f"t={at:>2}s  {audience:<20} {status:<9}  expires in {left:.0f}s")
Output
t= 0s  api://support-tools  new token  expires in 75s
t= 1s  api://billing        new token  expires in 74s
t= 5s  api://support-tools  reused     expires in 70s
t=20s  api://support-tools  new token  expires in 74s

At five seconds the cached token still had 70 seconds left, so it was reused. At twenty seconds it had 55 left, inside the 60-second window, so the identity fetched a fresh one. The billing token lives in its own cache entry. The cache is in memory and per process: a token is never written to disk, and two processes each fetch their own.

Warning

Keep token lifetimes well above 60 seconds. A token with less than a minute to live counts as stale the moment it arrives, so the identity fetches a new one on every call.


[07]

Swap in Entra ID, AWS, Google Cloud or SPIFFE

The local issuer was a stand-in. In production, the platform your agent runs on is the issuer, and the agent picks up its identity from the environment: a metadata service, or a token file the platform keeps fresh. You don't store a secret anywhere. Only the line that builds the identity changes; the server and build_agent stay exactly as they are.

Pythoncloud_identities.py
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
"""The same agent, backed by each cloud provider. Constructed here, not run against a cloud."""

from promptise.identity import AgentIdentity

identities = {
    "Microsoft Entra ID": AgentIdentity.from_entra(
        "support-bot", client_id="<managed-identity-client-id>", resource="api://support-tools"
    ),
    "AWS IAM": AgentIdentity.from_aws("support-bot", region="eu-central-1", audience="api://support-tools"),
    "Google Cloud": AgentIdentity.from_gcp("support-bot", audience="api://support-tools"),
    "SPIFFE / SPIRE": AgentIdentity.from_spiffe(
        "support-bot", token_file="/run/spiffe/jwt_svid.token", audience="api://support-tools"
    ),
    "Generic OIDC": AgentIdentity.from_oidc(
        "support-bot", issuer="https://token.actions.githubusercontent.com", token_env_var="AGENT_OIDC_TOKEN"
    ),
}

for platform, identity in identities.items():
    print(f"{platform:<20} {identity.credential_provider}")
Output
Microsoft Entra ID   entra-imds
AWS IAM              aws-sts
Google Cloud         gcp-metadata
SPIFFE / SPIRE       spiffe-file
Generic OIDC         oidc:https://token.actions.githubusercontent.com

This confirms each factory accepts these settings and picks the mode you'd expect off-platform. It doesn't prove the live round trip: none of these were run against a real cloud for this guide. Promptise's Security page is just as frank, and points to opt-in integration tests you can run inside each platform before you depend on one. On the server side, point JwksAuth.from_discovery at your platform's issuer URL instead of the local one.

Where the agent runs

Factory

Where the token comes from

Azure VM, Container Apps

from_entra()

Managed identity, through the instance metadata service

Azure Kubernetes Service

from_entra()

The file in AZURE_FEDERATED_TOKEN_FILE

AWS Lambda, EC2, ECS

from_aws()

STS GetWebIdentityToken, which needs boto3

AWS EKS

from_aws(mode="projected")

A projected token file, read from PROMPTISE_IDENTITY_TOKEN_FILE

Google Compute Engine, Cloud Run, GKE

from_gcp()

The metadata server

A SPIFFE or SPIRE mesh

from_spiffe()

The Workload API socket, or a spiffe-helper file

GitHub Actions, GitLab CI, any OIDC issuer

from_oidc()

A file, an environment variable or your own function

AgentIdentity.auto("support-bot") picks the provider for you from environment variables, and raises PlatformDetectionError when it finds none, for example on your laptop. Each provider has its own setup page under Credential providers.


[08]

Long-running agents and token expiry

build_agent fetches each server's token once, when it connects. The identity keeps refreshing its own cache, but the open connection keeps sending the token it started with. An agent that lives longer than its token will eventually be turned away. This run used 75-second tokens and waited past their expiry, plus the server's 60-second grace period:

Pythonlong_session.py
1
2
3
4
5
6
7
8
9
10
        start = time.time()
        await agent.ainvoke(QUESTION)
        await asyncio.sleep(140)  # longer than the token lifetime plus the server's 60s leeway
        print(f"--- {time.time() - start:.0f}s later")
        fresh = identity.get_credential("api://support-tools")
        print(f"The identity's own cached token now expires in {decode_jwt_expiry(fresh) - time.time():.0f}s")
        try:
            await asyncio.wait_for(agent.ainvoke(QUESTION), timeout=120)
        except TimeoutError:
            print("No answer after 120s: the tool call is still waiting.")
Output
→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 72}
--- 142s later
The identity's own cached token now expires in 75s
→ Invoking tool: whoami with {'payload': {}}
Server at http://127.0.0.1:8310/mcp answered HTTP 401 Unauthorized
No answer after 120s: the tool call is still waiting.

Two problems show here. The identity had a fresh token ready, but the connection never used it, so the server rejected the call. Worse, the rejected call didn't fail: it waited until the timeout gave up on it. Without asyncio.wait_for, an earlier run waited more than seven minutes before it was stopped by hand.

Until that changes, build a short-lived agent per job or per session, from the same identity. Each build connects with a token that has at least a minute left, and the shared identity means its cache is shared too:

Pythonagent_per_job.py
1
2
3
4
5
6
7
8
9
10
11
12
async def run_job(question: str) -> str:
    agent = await build_agent(
        model="openai:gpt-5-mini",
        servers={"support": HTTPServerSpec(url="http://127.0.0.1:8310/mcp", audience="api://support-tools")},
        identity=identity,  # the same identity, so its credential cache is shared
        trace_tools=True,
    )
    try:
        result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
        return result["messages"][-1].content
    finally:
        await agent.shutdown()
Output
→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 72}
>>> The whoami response shows token_expires_in_s = 72 (seconds).
--- 144s later
→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {"client_id": "support-bot", "issuer": "http://127.0.0.1:8311", "audience": "api://support-tools", "roles": ["support"], "token_expires_in_s": 73}
>>> I called whoami. The returned token_expires_in_s value is 73 — the token will expire in 73 seconds from the time of the call.

The second job got a new token and went through. Also put a timeout around ainvoke in anything long-running, and pick a token lifetime comfortably longer than your longest job.


[09]

The user's token stays in the agent

CallerContext has a bearer_token field, and it's tempting to expect the user's token to travel to the MCP server with each call. In 1.2.1 it doesn't; the field's own docstring calls MCP token forwarding "a planned enhancement". This check gives the agent no identity, passes a valid token as the caller's, and calls a server that checks tokens per tool:

Pythoncaller_token_check.py
        caller = CallerContext(user_id="maya@shop.example", bearer_token=token_from_idp("api://support-tools"))
        await agent.ainvoke({"messages": [{"role": "user", "content": "Call whoami once."}]}, caller=caller)
Output
→ Invoking tool: whoami with {'payload': {}}
✔ Tool result from whoami: {
  "error": {
    "code": "AUTHENTICATION_ERROR",
    "message": "Missing authentication token",
    "retryable": false
  }
}

So authenticate the agent with AgentIdentity, and keep user-level decisions where the user is known: in your app, in the agent's CallerContext for isolation and approvals, or as an explicit tool argument the server validates against its own records.


[10]

Honest limits

These are true of Promptise Foundry 1.2.1:

  • The token is fixed per connection, and a call made after it expires waits instead of failing, as shown above. Build an agent per job and wrap ainvoke in a timeout.

  • The user's token isn't forwarded to MCP servers, as shown above. A server sees the agent's identity only.

  • If the issuer is down when the agent connects, build_agent logs Agent identity 'support-bot' could not acquire a credential for MCP server audience 'api://support-tools' (…); connecting without it. and the server then refuses the connection with a 401. That's the safe direction, but watch for the warning in your logs.

  • An explicit `bearer_token` on an `HTTPServerSpec` wins over the identity for that server, so remove it when you switch to agent identity.

  • The server allows 60 seconds of clock skew. A token is accepted for up to a minute after its exp. JwksAuth(jwks_url=..., audience=..., leeway=...) lets you change that; JwksAuth.from_discovery doesn't take a leeway argument.

  • A new signing key is picked up automatically, but not instantly. When a token names a key id the server hasn't seen, it fetches the issuer's keys again, but only if its last fetch was more than 30 seconds ago. In this guide's runs, a token signed with a brand-new key was refused inside that window and accepted after it. Most issuers publish a new key before they start signing with it, which avoids the gap.

  • `subject()` and `idp_claims()` don't verify the token. They're for attribution inside your own process. Proof comes from the server checking the signature.

  • The cloud providers weren't run here. Run Promptise's opt-in integration tests on your platform before you rely on one.


[11]

Frequently asked questions

What is AI agent identity?

It's an identity for the agent itself, separate from the person using it: a name in your identity provider, plus a short-lived signed credential the agent presents to every service it calls. Services verify the credential, authorize the agent by name or role, and record which agent did what. In Promptise Foundry it's AgentIdentity, passed to build_agent(identity=...).

How do you authenticate an AI agent to an MCP server?

Have the agent present a signed JWT from your identity provider as a bearer token, and verify it on the server against the provider's public keys. In Promptise, give the agent an AgentIdentity and set audience on each HTTPServerSpec; on the server, add AuthMiddleware(JwksAuth.from_discovery(issuer=..., audience=...)) and require_auth=True. Then use guards like RequireClientId or HasRole to decide what each agent may do.

What is a non-human identity for AI agents?

A non-human identity is an identity for software rather than a person, like a service account, a managed identity or an IAM role. AI agents are a new kind of non-human identity: they act on their own, often all day, so they need credentials that are short-lived, scoped to one resource, and revocable in one place. Workload identity platforms such as Entra ID, AWS IAM, Google Cloud and SPIFFE already issue exactly that.

Do I still need user authentication if the agent has an identity?

Yes. The agent's identity proves which agent is calling; it says nothing about which person asked. Keep authenticating users in your app, pass who they are to the agent with CallerContext, and decide user-level permissions there, because in 1.2.1 the MCP server only sees the agent.

Does the agent need an API key or client secret?

Not on a cloud platform. The agent gets its token from the platform's metadata service or a token file the platform rotates, so there's no long-lived secret to store or leak. The server needs none either, because it verifies tokens with the issuer's public keys.


[12]

Where to go next

  • Agent Identity overview: local and verifiable identities, and where the identity shows up.

  • Agent Identity guide: an end-to-end scenario with delegation between agents.

  • Architecture and Security: credential lifecycle, per-resource credentials and the threat model.

  • Generic OIDC provider: from_oidc with files, functions and environment variables.

  • Authentication & Security: JwksAuth, JWTAuth, guards and ClientContext on the server.

  • Server configuration: every option on HTTPServerSpec, including audience.

  • How to Connect MCP Servers to Your AI Agent in Python: the agent and server basics this guide builds on.

  • OpenAPI to MCP: Turn Any REST API into an MCP Server: put an existing API behind an MCP server, then protect it the same way.

Learning paths

Want more structure? Paths put guides in order, like a short course.

See the paths →

Keep going.

Browse every guide by topic and level, or follow a learning path that puts them in order.

All guidesLearning paths