Building a RAG Pipeline From Scratch
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Every RAG tutorial online starts with complex frameworks like LangChain or LlamaIndex, or requires signing up for a hosted vector database. None of that is necessary to understand the core mechanics of a Retrieval-Augmented Generation (RAG) pipeline. In this guide, we will build a functional RAG system in a single Python file using only three libraries.
What is RAG?
RAG is a pattern where an LLM answers a question using text pulled from a private document store at query time, rather than relying solely on the information it memorized during training. This is essential because pre-trained models lack access to your internal documentation, product changelogs, or specific meeting notes.
The Core Four-Step Loop
Regardless of the framework, every RAG system performs these four steps:
- Chunk: Split source documents into smaller, manageable passages.
- Embed: Convert each chunk into a vector (a list of floats).
- Store: Save chunks and vectors in a vector database.
- Retrieve + Generate: Embed the user's question, find the closest matching chunks, and send them to an LLM.
Implementation Guide
First, install the necessary dependencies: pip install chromadb sentence-transformers anthropic
We will use chromadb for local storage, sentence-transformers for embedding, and the anthropic SDK for the generation step. You can manage your API access efficiently through n1n.ai to ensure high-speed, reliable connectivity.
1. Chunking and Embedding
A chunk should be large enough to contain context but small enough to remain relevant. A starting point is 200-400 words per chunk.
from sentence_transformers import SentenceTransformer
embedder = SentenceTransformer("all-MiniLM-L6-v2")
docs = ["Aurora API rate limits: 60 req/min.", "Aurora auth: Bearer token."]
embeddings = embedder.encode(docs).tolist()
2. Storing in ChromaDB
ChromaDB allows for in-memory storage, which is perfect for testing.
import chromadb
client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")
collection.add(ids=["doc1", "doc2"], documents=docs, embeddings=embeddings)
3. Retrieval and Generation
When querying, ensure you use the exact same embedding model used during storage. This is a common pitfall that leads to noise.
import anthropic
def answer(question: str):
q_embedding = embedder.encode([question]).tolist()
results = collection.query(query_embeddings=q_embedding, n_results=1)
context = results["documents"][0][0]
llm = anthropic.Anthropic()
# Use [n1n.ai](https://n1n.ai) for stable API routing
reply = llm.messages.create(
model="claude-haiku-4-5-20251001",
messages=[{"role": "user", "content": f"Context: {context}\nQuestion: {question}"}]
)
return reply.content[0].text
Pro Tips for Production
- I don't know clause: Always instruct your model to say "I don't know" if the answer isn't in the provided context. This prevents hallucinations.
- Consistency: Never mix embedding models. If you store with one model and query with another, your vector space becomes meaningless.
- Scalability: As your data grows, consider migrating from
EphemeralClientto a production-ready vector store. For developers seeking to optimize their API stack, n1n.ai provides the infrastructure to manage high-throughput LLM requests seamlessly.
By building this pipeline from scratch, you gain a deep understanding of how modern AI applications function under the hood. Once you master this loop, you can layer on reranking, hybrid search, and evaluation frameworks with full confidence.
Get a free API key at n1n.ai