Most RAG tutorials search your documents once, paste the top few chunks into the prompt, and hope they were the right ones. That breaks when a question needs two documents, or none. Agentic RAG turns it around: you give the AI agent your documents as a search tool, and the agent decides when to search, what to search for, and whether to search again. In this guide you'll index a small product handbook, hand it to a Python agent, read every search it makes, and teach it to say "I don't know" when the handbook has no answer. Every snippet ran against Promptise Foundry 1.2.1, and every output block is what it printed.
How do you build agentic RAG in Python?
Index your documents into a retrieval pipeline, wrap the pipeline as a tool, and give that tool to an agent. The model then uses it like any other tool: it writes a search query, reads the chunks that come back, and answers from them. With Promptise Foundry, that's RAGPipeline, rag_to_tool and build_agent:
from promptise import build_agent, rag_to_tool
pipeline = build_pipeline("handbook")
print(await pipeline.index())
search_handbook = rag_to_tool(
pipeline,
name="search_handbook",
description=(
"Search the Acme Inbox customer handbook: plans and pricing, refunds, "
"SSO, data retention, API rate limits, webhooks and support hours."
),
limit=4,
)
agent = await build_agent(
model="openai:gpt-5-mini",
servers={},
extra_tools=[search_handbook],
instructions=INSTRUCTIONS,
)build_pipeline is about 60 lines you'll write in step 2: a loader for your files, an embedder, and a vector store. The rest of this guide builds it and shows each piece working, including the part most tutorials skip: what the agent does when the answer isn't there.
[02]
How agentic RAG works
There are two phases. Indexing happens before any question: your files are split into chunks, each chunk is turned into a vector, and the vectors go into a store. At question time, the agent decides whether to search. When it does, the tool turns its query into a vector, finds the closest chunks, and hands them back with their sources.
Rendering diagram…
Why a tool instead of pasting documents into the prompt? The model writes its own search query, so it can rephrase a vague question. It can search twice when the first results don't fit, and you'll see it do exactly that below. And questions that need no documents cost no retrieval at all.
Promptise gives you the frame and two finished pieces. You write the loader and plug in an embedder:
Piece | In Promptise 1.2.1 | What you do |
|---|---|---|
Loader | DocumentLoader base class | Write load() for your source |
Chunker | RecursiveTextChunker | Use it, or subclass Chunker |
Embedder | Embedder base class | Wrap OpenAIEmbeddingProvider or LocalEmbeddingProvider |
Vector store | InMemoryVectorStore | Use it, or subclass VectorStore for your database |
Agent tool | rag_to_tool | Pass it to build_agent |
[03]
What you need
Python 3.10 or newer.
Promptise Foundry: pip install promptise. This guide used version 1.2.1.
An OpenAI API key in OPENAI_API_KEY. The agent uses it for the model and the pipeline uses it for embeddings, with text-embedding-3-small.
Optionally, pip install sentence-transformers to compute embeddings on your own machine instead.
[04]
Build a RAG agent over your docs, step by step
The example is a support assistant for Acme Inbox, a made-up shared inbox product. It answers customer questions from the product handbook.
Put your documents in a folder
Start with what you'd really have: a handful of Markdown pages, grouped by topic. The folder name will become each document's category, which you'll use for filtering later.
handbook/api/rate-limits.md
handbook/api/webhooks.md
handbook/billing/plans.md
handbook/billing/refunds.md
handbook/security/data-retention.md
handbook/security/sso.md
handbook/support/support-hours.mdEach file is a short page with a title and a few sections. Here's one:
# Refunds and cancellations
## Cancelling
Workspace owners can cancel at any time from Settings → Billing → Cancel plan. Your workspace keeps its paid features until the end of the period you already paid for, then moves to the Starter plan.
## Annual plans
If you cancel an annual plan within 30 days of the first annual invoice, we refund the full amount. After 30 days, annual plans are not refundable, but you keep access until the end of the year you paid for.
## Monthly plans
Monthly plans are not refunded. Cancel before the renewal date and you will not be charged again.
## How refunds are paid
Refunds go back to the original payment method. Card refunds usually appear within 5 to 10 business days, depending on the bank. Invoices paid by bank transfer are refunded by bank transfer within 10 business days.
To request a refund, email billing@acme-inbox.example with your workspace name and the invoice number.Write a loader and an embedder
Promptise doesn't guess where your documents live, so the loader is yours: a DocumentLoader with one load() method. The embedder wraps Promptise's OpenAI embedding client in the Embedder shape the pipeline expects:
from pathlib import Path
from promptise import (
Document,
DocumentLoader,
Embedder,
InMemoryVectorStore,
OpenAIEmbeddingProvider,
RAGPipeline,
RecursiveTextChunker,
)
class MarkdownFolderLoader(DocumentLoader):
"""Load every Markdown file under a folder as one Document."""
def __init__(self, root: str) -> None:
self.root = Path(root)
async def load(self) -> list[Document]:
docs = []
for path in sorted(self.root.rglob("*.md")):
text = path.read_text(encoding="utf-8")
source = path.relative_to(self.root).as_posix()
docs.append(
Document(
id=source,
text=text,
metadata={
"source": source,
"title": text.splitlines()[0].lstrip("# ").strip(),
"category": path.parent.name,
},
)
)
return docs
class OpenAIEmbedder(Embedder):
"""Promptise's OpenAI embedding client, shaped as a RAG Embedder."""
def __init__(self, model: str = "text-embedding-3-small") -> None:
self._provider = OpenAIEmbeddingProvider(model=model, api_key="${OPENAI_API_KEY}")
async def embed(self, texts: list[str]) -> list[list[float]]:
return await self._provider.embed(texts)
@property
def dimension(self) -> int:
return 1536
def build_pipeline(root: str = "handbook") -> RAGPipeline:
return RAGPipeline(
loader=MarkdownFolderLoader(root),
chunker=RecursiveTextChunker(chunk_size=500, overlap=50),
embedder=OpenAIEmbedder(),
store=InMemoryVectorStore(dimension=1536),
)What each choice buys you:
The document ID is the file path. Chunk IDs are built from it (billing/refunds.md:chunk-0), so the same file always maps to the same IDs. That's what makes updates and deletes possible later.
Metadata travels with every chunk. The chunker copies source, title and category onto each chunk. rag_to_tool prints the source and title above every result, which is how the agent can cite them.
The key stays out of your code. "${OPENAI_API_KEY}" is read from the environment when the embedder is called.
`chunk_size=500, overlap=50` are the chunker's defaults, written out so you can see them. InMemoryVectorStore(dimension=1536) rejects any vector of the wrong length, and the rejection shows up in the index report's errors. That catches a mismatched embedding model early.
Index the handbook and test a search
Check retrieval on its own before an agent is involved. If the right chunks don't come back here, no prompt will fix it later.
import asyncio
from handbook_rag import build_pipeline
async def main():
pipeline = build_pipeline("handbook")
report = await pipeline.index()
print(report)
hits = await pipeline.retrieve("Can we use Okta to sign in?", limit=3)
for hit in hits:
print(f"{hit.score:.2f} {hit.chunk.id} {hit.text[:60]!r}")
asyncio.run(main())IndexReport(documents_loaded=7, chunks_created=18, chunks_stored=18, errors=[], duration_seconds=0.8770301250006014)
0.76 security/sso.md:chunk-0 '# Single sign-on (SSO)\n\nSAML single sign-on is available on '
0.74 security/sso.md:chunk-1 'Azure AD) and Google Workspace.\n\n## Setting it up\n\n1. Go to '
0.72 security/sso.md:chunk-2 'Test connection, then Enable.\n\n## Enforcing SSO\n\nOnce SSO wo'The IndexReport is your health check: 7 files became 18 chunks, and all 18 were stored. If a file fails to chunk or a batch fails to embed, index() keeps going and lists the failure in errors, so check it rather than assuming success.
The question never says "SSO", yet all three hits come from the SSO page. That's semantic search working. Each result is a RetrievalResult with the chunk, its metadata and a score between 0 and 1. Keep one detail in mind: InMemoryVectorStore maps cosine similarity onto that range, so even unrelated text scores around 0.6. You'll see why that matters in a moment.
Give the agent the search tool
Now hand the pipeline to an agent. rag_to_tool turns it into a tool with a query argument and a limit argument, and extra_tools gives it to the agent:
import asyncio
from langchain_core.messages import AIMessage, ToolMessage
from handbook_rag import build_pipeline
from promptise import build_agent, rag_to_tool
INSTRUCTIONS = """You are a support assistant for Acme Inbox.
Answer only from the handbook. Call search_handbook before you answer,
and use only what the results say. End with the sources you used,
for example (Source: billing/refunds.md).
If the handbook doesn't answer the question, say you don't know and
suggest emailing support@acme-inbox.example. Never guess prices, dates or features."""
QUESTIONS = [
"We're on Team with 15 people. Can we use Okta for SSO, and what would it cost us per month?",
"A customer cancelled their annual plan after 3 weeks. Do they get their money back, and how long does it take?",
"Does Acme Inbox integrate with Salesforce?",
]
def print_trace(result):
"""Show each tool call the model made and what came back."""
for message in result["messages"]:
if isinstance(message, AIMessage):
for call in message.tool_calls:
print(f"→ {call['name']}({call['args']})")
elif isinstance(message, ToolMessage):
print(f"✔ Result:\n{message.content}")
async def main():
pipeline = build_pipeline("handbook")
print(await pipeline.index())
search_handbook = rag_to_tool(
pipeline,
name="search_handbook",
description=(
"Search the Acme Inbox customer handbook: plans and pricing, refunds, "
"SSO, data retention, API rate limits, webhooks and support hours."
),
limit=4,
)
agent = await build_agent(
model="openai:gpt-5-mini",
servers={},
extra_tools=[search_handbook],
instructions=INSTRUCTIONS,
)
try:
for question in QUESTIONS:
print(f"\n### {question}")
result = await agent.ainvoke({"messages": [{"role": "user", "content": question}]})
print_trace(result)
print(">>>", result["messages"][-1].content)
finally:
await agent.shutdown()
asyncio.run(main())A few things to know:
The description is how the model decides to search. Say what's in the knowledge base, in the words your users will use. A vague description gets a tool the model forgets to call.
`servers={}` is fine. The agent needs no MCP servers when its only tool is the search tool. You can add servers next to it at any time.
`limit=4` is a default, not a cap. The tool exposes limit to the model, and the model may ask for more. In the run below it asked for 5, 6 and 10.
The trace comes from the messages. result["messages"] holds every tool call and every tool result, so print_trace shows exactly what the model searched for and read.
Ask questions that need the documents
python agent.pyHere's the first question, with its search and the chunks it read:
IndexReport(documents_loaded=7, chunks_created=18, chunks_stored=18, errors=[], duration_seconds=2.4973907919993508)
### We're on Team with 15 people. Can we use Okta for SSO, and what would it cost us per month?
→ search_handbook({'query': 'Okta SSO pricing Team 15 people per month', 'limit': 5})
✔ Result:
Found 5 result(s):
### Result 1 (relevance: 0.78) — Plans and pricing
*Source: `billing/plans.md`*
and up to 20 shared inboxes.
## Business
$24 per seat per month, billed monthly, or $240 per seat per year when billed annually. Includes everything in Team, plus SAML single sign-on (SSO), SCIM user provisioning, the audit log and a 4-hour first response time from support.
## Changing seats or plans
### Result 2 (relevance: 0.75) — Single sign-on (SSO)
*Source: `security/sso.md`*
# Single sign-on (SSO)
SAML single sign-on is available on the Business plan. It is not available on Starter or Team. Workspaces on Team must upgrade to Business before they can enable SSO.
…
>>> Short answer: Yes — but only on the Business plan. Okta is a supported SAML provider, and SSO requires upgrading from Team to Business.
Details:
- SSO (SAML) is available only on the Business plan; Team workspaces must upgrade to Business to enable SSO. (Source: security/sso.md)
- We support Okta (Acme Inbox tests and documents setup for Okta). (Source: security/sso.md)
- Business plan pricing: $24 per seat per month (monthly billing). For 15 seats that is 15 × $24 = $360 per month. (Source: billing/plans.md)
…The answer needed two pages: SSO says Team must upgrade, and pricing says what Business costs. The model wrote one query that pulled both, then did the multiplication itself and cited each fact. Notice the first result starts mid-sentence, "and up to 20 shared inboxes." That's the chunk overlap: each chunk repeats the tail of the one before it. The title in the result header tells the model which page it came from anyway.
The second question, with the retrieved chunks trimmed:
### A customer cancelled their annual plan after 3 weeks. Do they get their money back, and how long does it take?
→ search_handbook({'query': 'refund annual plan cancelled after 3 weeks money back how long does it take refund policy annual cancellation 3 weeks', 'limit': 6})
✔ Result:
Found 6 result(s):
### Result 1 (relevance: 0.75) — Refunds and cancellations
*Source: `billing/refunds.md`*
…
>>> Yes. Canceling an annual plan within 30 days of the first annual invoice qualifies for a full refund — 3 weeks is within that window.
Refund timing:
- Refunds go back to the original payment method.
- Card refunds usually appear within 5–10 business days (depending on the bank).
- Invoices paid by bank transfer are refunded by bank transfer within 10 business days.
To request a refund, email billing@acme-inbox.example with your workspace name and the invoice number.
(Source: billing/refunds.md)The model compared "3 weeks" to the 30-day rule, which no keyword search would do for it. Your wording will differ from run to run, because the model writes each answer fresh.
[05]
When the answer isn't in your documents
This is where RAG agents get into trouble. The handbook says nothing about Salesforce. Here's the same agent with a one-line instruction instead of the grounding rules, from agent_naive.py:
INSTRUCTIONS = "You are a support assistant for Acme Inbox."### Does Acme Inbox integrate with Salesforce?
→ search_handbook({'query': 'Salesforce integration', 'limit': 4})
✔ Result:
Found 4 result(s):
### Result 1 (relevance: 0.69) — Single sign-on (SSO)
*Source: `security/sso.md`*
…
→ search_handbook({'query': 'Salesforce Acme Inbox CRM integration Zapier webhooks API Salesforce', 'limit': 4})
…
>>> Short answer: No — there’s no built-in, native Salesforce app in Acme Inbox.
You can integrate though. Options:
…
If it helps, I can:
- Show a sample webhook payload and the signature verification steps.
- Draft a simple middleware example (Node/Python) that turns an Acme webhook into a Salesforce Case.
- Suggest a Zapier/Make recipe if you want a no-code route.
…Two things went wrong, and both are normal. First, the search returned something anyway: SSO and pricing chunks, with relevance around 0.7. A vector store always returns the closest chunks, even when none of them are relevant. Second, the model filled the gap. "There's no built-in, native Salesforce app" is a claim the handbook never makes, and the offer of sample payloads and Zapier recipes is help it can't back up.
The grounding rules in agent.py change that. Here is the third question from the run above:
### Does Acme Inbox integrate with Salesforce?
→ search_handbook({'query': 'Salesforce integration Acme Inbox', 'limit': 4})
✔ Result:
Found 4 result(s):
### Result 1 (relevance: 0.77) — Single sign-on (SSO)
*Source: `security/sso.md`*
…
→ search_handbook({'query': 'Salesforce', 'limit': 10})
✔ Result:
Found 10 result(s):
…
>>> I don't know — the handbook has no mention of a native Salesforce integration. If you need to connect Acme Inbox to Salesforce, the handbook does note there is a REST API and webhooks (available on Team and Business) you could use to build an integration. For a definitive answer or roadmap, please email support@acme-inbox.example.
(Sources: api/webhooks.md, api/rate-limits.md, billing/plans.md, support/support-hours.md)That's agentic RAG doing its job. The first search came back with nothing useful, so the agent searched again with a shorter query and a bigger limit. When that failed too, it said so, pointed to what the handbook does cover, and sent the customer to a person.
Why not filter out weak results by score instead? In this run, the best unrelated chunk for Salesforce scored 0.77, higher than the best chunk for the refund question (0.75). A cut-off would have dropped good answers before it dropped bad ones. rag_to_tool has no minimum score in 1.2.1 anyway, so the instructions do the work.
[06]
Choosing a chunk size
Chunk size decides what one search result contains. Too small and a result lacks context; too large and the answer is buried in text the model has to read anyway. Here's the same question against three sizes:
import asyncio
from handbook_rag import MarkdownFolderLoader, OpenAIEmbedder
from promptise import InMemoryVectorStore, RAGPipeline, RecursiveTextChunker
QUESTION = "How long do card refunds take?"
async def main():
for size in (200, 500, 1500):
pipeline = RAGPipeline(
loader=MarkdownFolderLoader("handbook"),
chunker=RecursiveTextChunker(chunk_size=size, overlap=size // 10),
embedder=OpenAIEmbedder(),
store=InMemoryVectorStore(),
)
report = await pipeline.index()
top = (await pipeline.retrieve(QUESTION, limit=1))[0]
print(f"chunk_size={size}: {report.chunks_created} chunks, top hit {top.chunk.id} ({len(top.text)} chars)")
print(" ", top.text.replace("\n", " ")[:110], "…")
asyncio.run(main())chunk_size=200: 74 chunks, top hit billing/refunds.md:chunk-10 (145 chars)
refunds are paid Refunds go back to the original payment method. Card refunds usually appear within 5 to 10 b …
chunk_size=500: 18 chunks, top hit billing/refunds.md:chunk-1 (491 chars)
end of the year you paid for. ## Monthly plans Monthly plans are not refunded. Cancel before the renewal dat …
chunk_size=1500: 7 chunks, top hit billing/refunds.md:chunk-0 (931 chars)
# Refunds and cancellations ## Cancelling Workspace owners can cancel at any time from Settings → Billing → …All three found the right page. At 200 characters the hit is exactly the right paragraph, but the heading was cut to "refunds are paid", and 74 chunks is a lot for seven short pages. At 1,500 each page is one chunk, so every result is a whole page. For pages like these, 500 is a good middle: one or two sections per result. chunk_size counts characters, not tokens, and the chunker splits on blank lines first, then line breaks, then sentences, then words.
Very small sizes also produce stray fragments. At chunk_size=200, the refunds page yields chunks that are only a heading, such as Cancelling and Annual plans. If your documents have clear headings, a custom Chunker that splits on them is often better than any size; the RAG docs show the shape.
[07]
Filtering by metadata
Some questions belong to one part of your documents. "What happens after 30 days?" means one thing for refunds and another for deleted workspaces. retrieve() takes a filter that matches chunk metadata:
import asyncio
from handbook_rag import build_pipeline
QUESTION = "What happens after 30 days?"
async def main():
pipeline = build_pipeline("handbook")
await pipeline.index()
for scope in (None, {"category": "billing"}, {"category": "security"}):
hits = await pipeline.retrieve(QUESTION, limit=2, filter=scope)
print(f"filter={scope}")
for hit in hits:
print(f" {hit.score:.2f} {hit.metadata['source']}")
asyncio.run(main())filter=None
0.69 security/data-retention.md
0.68 security/data-retention.md
filter={'category': 'billing'}
0.64 billing/refunds.md
0.64 billing/refunds.md
filter={'category': 'security'}
0.69 security/data-retention.md
0.68 security/data-retention.mdWithout a filter, data retention wins. Scoped to billing, the refund policy comes back instead. InMemoryVectorStore filters by exact match on every key you pass; there are no ranges or "any of" operators.
rag_to_tool doesn't pass a filter in 1.2.1, so the agent can't scope its own searches. If you need that, build one pipeline per scope and give the agent one tool for each, with a description that says what's inside. The RAG docs call this multi-store composition.
[08]
Keeping the index fresh
Documents change. The pipeline won't notice on its own, and one detail catches people out: index() adds and overwrites chunks, but never removes old ones. Re-index a page that got shorter and its old tail stays in the store. Here the refunds page is replaced by a two-sentence version, without deleting the old one first, from stale.py:
# Re-index the new, shorter version WITHOUT deleting the old one first.
await pipeline.index(documents=[Document(id="billing/refunds.md", text=NEW)])old version: 2 chunks
new version: 2 chunks
0.78 billing/refunds.md:chunk-0: 'Card refunds appear within 3 to 5 business days.'…
0.72 billing/refunds.md:chunk-1: 'Card refunds usually appear within 5 to 10 business days, de'…The new page has one chunk, but the store still holds two. The agent would now read both "3 to 5" and "5 to 10" business days. The fix is to call delete_document() before re-indexing a changed page. This sync function does that, skips pages that didn't change, and removes pages that were deleted:
import asyncio
import shutil
from pathlib import Path
from handbook_rag import MarkdownFolderLoader, OpenAIEmbedder
from promptise import InMemoryVectorStore, RAGPipeline, RecursiveTextChunker, content_hash
async def sync(pipeline: RAGPipeline, loader: MarkdownFolderLoader, seen: dict[str, str]) -> str:
"""Re-index new and changed files, and remove deleted ones."""
docs = await loader.load()
changed = [d for d in docs if seen.get(d.id) != content_hash(d.text)]
removed = set(seen) - {d.id for d in docs}
for doc_id in [d.id for d in changed] + list(removed):
await pipeline.delete_document(doc_id) # drop the old chunks first
seen.pop(doc_id, None)
report = await pipeline.index(documents=changed)
seen.update({d.id: content_hash(d.text) for d in changed})
return f"{len(changed)} changed, {len(removed)} removed, {report.chunks_stored} chunks embedded"
async def main():
live = Path("handbook-live")
shutil.rmtree(live, ignore_errors=True)
shutil.copytree("handbook", live)
loader = MarkdownFolderLoader(str(live))
store = InMemoryVectorStore()
pipeline = RAGPipeline(
chunker=RecursiveTextChunker(chunk_size=500, overlap=50),
embedder=OpenAIEmbedder(),
store=store,
)
seen: dict[str, str] = {}
print("first sync: ", await sync(pipeline, loader, seen), "| store:", await store.count())
print("no changes: ", await sync(pipeline, loader, seen), "| store:", await store.count())
refunds = live / "billing" / "refunds.md"
refunds.write_text(
"# Refunds and cancellations\n\n"
"Annual plans cancelled within 14 days of the first annual invoice are refunded in full. "
"Card refunds appear within 3 to 5 business days.\n"
)
(live / "api" / "webhooks.md").unlink()
print("after edits: ", await sync(pipeline, loader, seen), "| store:", await store.count())
for hit in await pipeline.retrieve("How long do card refunds take?", limit=2):
print(f" {hit.score:.2f} {hit.chunk.id}: …{hit.text[-60:]!r}")
asyncio.run(main())first sync: 7 changed, 0 removed, 18 chunks embedded | store: 18
no changes: 0 changed, 0 removed, 0 chunks embedded | store: 18
after edits: 1 changed, 1 removed, 1 chunks embedded | store: 15
0.78 billing/refunds.md:chunk-0: …'ed in full. Card refunds appear within 3 to 5 business days.'
0.61 billing/plans.md:chunk-2: …'rged pro rata. Downgrading takes effect at the next renewal.'A sync with no changes embeds nothing, so you can run it on a schedule or after every deploy of your docs. After the edits, the store went from 18 chunks to 15: the refunds page shrank to one chunk, and the two webhook chunks are gone. The only refund answer left is the new one. content_hash gives a short, stable hash of a text, so an unchanged page is never embedded twice.
InMemoryVectorStore lives in your process, so a restart starts from an empty index and the first sync embeds everything again. For a handbook this size that took a few seconds in these runs. For a larger corpus, keep the vectors in a database by subclassing VectorStore; see the honest limits below.
[09]
Run embeddings on your own machine
If your documents can't leave your network, compute embeddings locally. LocalEmbeddingProvider runs a sentence-transformers model, all-MiniLM-L6-v2 by default, and wraps the same way as the OpenAI one. After pip install sentence-transformers:
class LocalEmbedder(Embedder):
"""Embeddings computed on your machine with sentence-transformers."""
def __init__(self, model: str = "all-MiniLM-L6-v2") -> None:
self._provider = LocalEmbeddingProvider(model=model)
async def embed(self, texts: list[str]) -> list[list[float]]:
return await self._provider.embed(texts)
@property
def dimension(self) -> int:
return 384Use LocalEmbedder() and InMemoryVectorStore(dimension=384) in the pipeline, and nothing else changes:
IndexReport(documents_loaded=7, chunks_created=18, chunks_stored=18, errors=[], duration_seconds=129.50240979099908)
0.78 security/sso.md:chunk-0
0.68 security/sso.md:chunk-2
0.68 security/sso.md:chunk-1Same question, same top hit. The long duration is the first embedding call importing sentence-transformers and loading the model, which happens once per process. Vectors from different models can't be compared, so when you switch embedders, re-index everything. The model you pass to build_agent is a separate choice: Models & Providers lists local options such as Ollama.
[10]
Honest limits
Promptise's RAG module is a small, clear frame, not a full ingestion suite. In 1.2.1:
One vector store ships: `InMemoryVectorStore`. It doesn't persist and searches every chunk on each query, which suits small corpora and tests. For Chroma, Pinecone, Qdrant or pgvector, subclass VectorStore and implement add, search and delete, plus delete_by_document so delete_document() works. ChromaProvider in Promptise is a backend for agent memory, not a RAG store.
No ready-made loaders. DocumentLoader is a base class. PDFs, web pages or Confluence need your own load(), as in step 2.
The embedding providers need a wrapper. OpenAIEmbeddingProvider and LocalEmbeddingProvider come from Promptise's semantic cache and don't have the embed_one method the pipeline calls.
Re-indexing doesn't replace old chunks. Delete a changed document before you index it again, as sync does.
`rag_to_tool` is simple. No metadata filter, no minimum score, and the model can raise limit without a ceiling.
[11]
Frequently asked questions
What is agentic RAG?
Agentic RAG is retrieval-augmented generation where the agent controls the retrieval. Instead of your code searching once and pasting results into the prompt, the agent gets search as a tool. It decides whether to search, writes the query, and can search again with a different query when the first results don't answer the question.
What's the difference between RAG and agentic RAG?
Traditional RAG runs a fixed step before the model: embed the user's message, fetch the top chunks, add them to the prompt. Agentic RAG lets the model choose. That costs an extra model turn per search, but it handles vague questions, questions that span several documents, and questions that need no documents at all. In this guide, the agent searched twice before deciding the handbook had no answer.
Can I use Chroma, Pinecone or pgvector with Promptise?
Yes, by subclassing VectorStore: implement add, search and delete against your database, and return RetrievalResult objects with scores between 0 and 1. Promptise 1.2.1 ships no database store of its own. The rest of the pipeline and rag_to_tool work unchanged with your store.
Is RAG the same as agent memory?
No. RAG searches documents you indexed ahead of time, and the agent calls it when it needs to. Memory holds what the agent learned about a user in earlier conversations, and Promptise adds it to the context automatically. Many agents use both. See Memory in the docs.
Can I build a RAG agent without OpenAI?
Yes. Use LocalEmbeddingProvider for embeddings, as shown above, and any model with tool calling for the agent, by changing the model string. The pipeline and the tool don't depend on the provider.
[12]
Where to go next
RAG: the pipeline, the base classes and production patterns.
RAG API reference: every class and parameter in promptise.rag.
Semantic cache: where OpenAIEmbeddingProvider and LocalEmbeddingProvider come from.
Memory: give your agent what it learned about each user.
Building agents: every option on build_agent, including extra_tools.
How to Connect MCP Servers to Your AI Agent in Python: give the same agent tools that act, next to the search tool.