Build a system that answers questions about PDF documents by converting text into numerical vectors and retrieving relevant context for an LLM.
What it is
A PDF Q&A bot uses Retrieval-Augmented Generation (RAG). It splits PDFs into chunks, converts them into embeddings (numerical representations of meaning), stores these in a vector database, and retrieves the most similar chunks when a user asks a question. The retrieved text is then passed to a Large Language Model (LLM) to generate an answer grounded in the document.
Key terms: Embeddings (vector representations), Vector Store (database for similarity search), Chunking (splitting text), Prompt Engineering (structuring input for the LLM).
Why it matters
- Accuracy: Reduces hallucinations by grounding answers in specific source text.
- Efficiency: Avoids sending entire large documents to the LLM, saving tokens and cost.
- Scalability: Allows querying massive document repositories instantly.
- Privacy: Can be run locally or on private infrastructure without exposing data to public APIs.
Syntax or steps
- Ingest: Load PDF and extract text.
- Split: Break text into manageable chunks (e.g., 500-1000 characters).
- Embed: Convert each chunk into a vector using an embedding model.
- Store: Save vectors and original text in a vector store.
- Retrieve: Embed the user's query and find the top-k most similar chunks.
- Generate: Combine retrieved chunks with the query into a prompt for the LLM.
Example
import fitz # PyMuPDF
from langchain.text_splitter import CharacterTextSplitter
from langchain_community.vectorstores import FAISS
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
# 1. Load PDF and Extract Text
doc = fitz.open("example.pdf")
text = ""
for page in doc:
text += page.get_text()
# 2. Split Text into Chunks
splitter = CharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_text(text)
# 3. Create Embeddings and Vector Store
embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")
db = FAISS.from_texts(chunks, embeddings)
# 4. Setup LLM and Prompt
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
prompt = ChatPromptTemplate.from_template(
"Answer the question based only on the following context:\n"
"{context}\n\nQuestion: {question}"
)
# 5. Retrieve and Generate
def ask_question(query):
docs = db.similarity_search(query, k=3)
context = "\n".join([d.page_content for d in docs])
chain = prompt | llm
response = chain.invoke({"context": context, "question": query})
return response.content
print(ask_question("What are the main findings?"))
This code extracts text from a PDF, creates a local vector index using FAISS, and uses OpenAI’s GPT model to answer questions based on the retrieved context. Note that API keys must be set in environment variables for `ChatOpenAI` to work.
Common mistakes
- Poor Chunking: Chunks too small lose context; too large dilute relevance. Use overlap to maintain continuity.
- Ignoring Metadata: Failing to track which page or section a chunk came from makes citations impossible.
- Static Retrieval: Using only simple similarity search can miss synonyms. Consider hybrid search (keyword + semantic).
- Over-reliance on LLM: If the retrieval fails, the LLM will guess. Always verify if the context contains the answer before generating.
When to use it
| Approach | Best For | Limitations |
|---|---|---|
| RAG (PDF Q&A) | Large, static, or frequently updated document sets where accuracy and citation are critical. | Complex setup; requires maintenance of vector databases. |
| Fine-Tuning | Changing the model's style, tone, or domain-specific knowledge permanently. | Expensive; does not easily update with new facts; high risk of hallucination if data is outdated. |
| Long Context Window | Small documents that fit entirely within the LLM's token limit. | Costly per query; performance degrades as context length increases ("lost in the middle"). |
Practice
Guided Exercise: Modify the example to print the source page number for each retrieved chunk. Hint: Store metadata during the `FAISS.from_texts` step or use `similarity_search_with_score`.
Challenge: Implement a "no answer" guardrail. If the similarity score of the top result is below a threshold, return "I cannot find this information in the document." instead of asking the LLM.
Quick check
Q: Why do we split PDFs into chunks before embedding?
A: Embedding models have token limits, and smaller chunks allow for more precise semantic matching between specific queries and relevant text segments.
Summary
PDF Q&A bots leverage RAG to combine the reasoning power of LLMs with the factual grounding of external documents. By structuring data into searchable vectors, you enable accurate, scalable, and cost-effective document interaction.