Simple RAG — Ask Your PDF

A simple Retrieval-Augmented Generation application that lets users upload a PDF and ask questions using embeddings, vector search, and an LLM.

PythonOpenAIFAISSRAGEmbeddings

Designed, built, and operated by Amal Chaitanya

Problem

Most RAG tutorials hide the retrieval step behind a framework. It is hard to see where chunking ends and generation begins, which makes debugging bad answers guesswork.

Why I built it

I built Simple RAG to expose every stage — ingestion, chunking, embedding, FAISS lookup, prompt assembly — so each failure mode (bad chunks, weak retrieval, poor prompting) is observable and fixable.

How it works

  • PyPDF extracts raw text and preserves page numbers for citations.
  • A recursive splitter creates ~500-token overlapping chunks; overlap keeps tables and paragraph boundaries intact.
  • Each chunk is embedded once and stored in an in-memory FAISS index for semantic retrieval.
  • At query time, top-k chunks are retrieved by cosine similarity and assembled into a grounded prompt.
  • The LLM answers only from retrieved context and returns chunk sources with page references.

Technology stack

  • Python + FastAPI for the API layer
  • FAISS for in-memory vector search — no managed DB required for the demo path
  • OpenAI embeddings + chat completions, swappable for local models
  • Minimal React UI showing retrieved chunks alongside the final answer

Challenges

  • Chunk size vs. recall: 200-token chunks were precise but lost context; 1000-token chunks diluted similarity scores.
  • Scanned PDFs needed OCR fallback — text extraction silently returned empty pages.
  • Evaluation: added a small golden-question set to compare chunking strategies objectively.

What I learned

  • Chunking matters more than model choice for answer quality.
  • Returning sources changes user trust — people verify instead of blindly accepting.
  • A minimal RAG is the right foundation before adding rerankers, HyDE, or agents.