NEWn1n v2.0.1 is live! Enterprise Unified LLM API Gateway with 500+ AI Models, up to 90% off, Try now

Building a RAG Pipeline From Scratch

Authors
  • avatar
    Name
    Nino
    Occupation
    Senior Tech Editor

Every RAG tutorial online starts with complex frameworks like LangChain or LlamaIndex, or requires signing up for a hosted vector database. None of that is necessary to understand the core mechanics of a Retrieval-Augmented Generation (RAG) pipeline. In this guide, we will build a functional RAG system in a single Python file using only three libraries.

What is RAG?

RAG is a pattern where an LLM answers a question using text pulled from a private document store at query time, rather than relying solely on the information it memorized during training. This is essential because pre-trained models lack access to your internal documentation, product changelogs, or specific meeting notes.

The Core Four-Step Loop

Regardless of the framework, every RAG system performs these four steps:

  1. Chunk: Split source documents into smaller, manageable passages.
  2. Embed: Convert each chunk into a vector (a list of floats).
  3. Store: Save chunks and vectors in a vector database.
  4. Retrieve + Generate: Embed the user's question, find the closest matching chunks, and send them to an LLM.

Implementation Guide

First, install the necessary dependencies: pip install chromadb sentence-transformers anthropic

We will use chromadb for local storage, sentence-transformers for embedding, and the anthropic SDK for the generation step. You can manage your API access efficiently through n1n.ai to ensure high-speed, reliable connectivity.

1. Chunking and Embedding

A chunk should be large enough to contain context but small enough to remain relevant. A starting point is 200-400 words per chunk.

from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer("all-MiniLM-L6-v2")
docs = ["Aurora API rate limits: 60 req/min.", "Aurora auth: Bearer token."]
embeddings = embedder.encode(docs).tolist()

2. Storing in ChromaDB

ChromaDB allows for in-memory storage, which is perfect for testing.

import chromadb

client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")
collection.add(ids=["doc1", "doc2"], documents=docs, embeddings=embeddings)

3. Retrieval and Generation

When querying, ensure you use the exact same embedding model used during storage. This is a common pitfall that leads to noise.

import anthropic

def answer(question: str):
    q_embedding = embedder.encode([question]).tolist()
    results = collection.query(query_embeddings=q_embedding, n_results=1)
    context = results["documents"][0][0]
    
    llm = anthropic.Anthropic()
    # Use [n1n.ai](https://n1n.ai) for stable API routing
    reply = llm.messages.create(
        model="claude-haiku-4-5-20251001",
        messages=[{"role": "user", "content": f"Context: {context}\nQuestion: {question}"}]
    )
    return reply.content[0].text

Pro Tips for Production

  • I don't know clause: Always instruct your model to say "I don't know" if the answer isn't in the provided context. This prevents hallucinations.
  • Consistency: Never mix embedding models. If you store with one model and query with another, your vector space becomes meaningless.
  • Scalability: As your data grows, consider migrating from EphemeralClient to a production-ready vector store. For developers seeking to optimize their API stack, n1n.ai provides the infrastructure to manage high-throughput LLM requests seamlessly.

By building this pipeline from scratch, you gain a deep understanding of how modern AI applications function under the hood. Once you master this loop, you can layer on reranking, hybrid search, and evaluation frameworks with full confidence.

Get a free API key at n1n.ai