Simple RAG — Ask Your PDF
A simple Retrieval-Augmented Generation application that lets users upload a PDF and ask questions using embeddings, vector search, and an LLM.
PythonOpenAIFAISSRAGEmbeddings
Designed, built, and operated by Amal Chaitanya
Problem
Most RAG tutorials hide the retrieval step behind a framework. It is hard to see where chunking ends and generation begins, which makes debugging bad answers guesswork.
Why I built it
I built Simple RAG to expose every stage — ingestion, chunking, embedding, FAISS lookup, prompt assembly — so each failure mode (bad chunks, weak retrieval, poor prompting) is observable and fixable.
How it works
- PyPDF extracts raw text and preserves page numbers for citations.
- A recursive splitter creates ~500-token overlapping chunks; overlap keeps tables and paragraph boundaries intact.
- Each chunk is embedded once and stored in an in-memory FAISS index for semantic retrieval.
- At query time, top-k chunks are retrieved by cosine similarity and assembled into a grounded prompt.
- The LLM answers only from retrieved context and returns chunk sources with page references.
Technology stack
- Python + FastAPI for the API layer
- FAISS for in-memory vector search — no managed DB required for the demo path
- OpenAI embeddings + chat completions, swappable for local models
- Minimal React UI showing retrieved chunks alongside the final answer
Challenges
- Chunk size vs. recall: 200-token chunks were precise but lost context; 1000-token chunks diluted similarity scores.
- Scanned PDFs needed OCR fallback — text extraction silently returned empty pages.
- Evaluation: added a small golden-question set to compare chunking strategies objectively.
What I learned
- Chunking matters more than model choice for answer quality.
- Returning sources changes user trust — people verify instead of blindly accepting.
- A minimal RAG is the right foundation before adding rerankers, HyDE, or agents.