Back to Data Science Notes
Topic #91

PDF Q&A Bot with an LLM

Build a system that answers questions about PDF documents by converting text into numerical vectors and retrieving relevant context for an LLM.

What it is

A PDF Q&A bot uses Retrieval-Augmented Generation (RAG). It splits PDFs into chunks, converts them into embeddings (numerical representations of meaning), stores these in a vector database, and retrieves the most similar chunks when a user asks a question. The retrieved text is then passed to a Large Language Model (LLM) to generate an answer grounded in the document.

Key terms: Embeddings (vector representations), Vector Store (database for similarity search), Chunking (splitting text), Prompt Engineering (structuring input for the LLM).

Why it matters

  • Accuracy: Reduces hallucinations by grounding answers in specific source text.
  • Efficiency: Avoids sending entire large documents to the LLM, saving tokens and cost.
  • Scalability: Allows querying massive document repositories instantly.
  • Privacy: Can be run locally or on private infrastructure without exposing data to public APIs.

Syntax or steps

  1. Ingest: Load PDF and extract text.
  2. Split: Break text into manageable chunks (e.g., 500-1000 characters).
  3. Embed: Convert each chunk into a vector using an embedding model.
  4. Store: Save vectors and original text in a vector store.
  5. Retrieve: Embed the user's query and find the top-k most similar chunks.
  6. Generate: Combine retrieved chunks with the query into a prompt for the LLM.

Example

import fitz  # PyMuPDF
from langchain.text_splitter import CharacterTextSplitter
from langchain_community.vectorstores import FAISS
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI

# 1. Load PDF and Extract Text
doc = fitz.open("example.pdf")
text = ""
for page in doc:
    text += page.get_text()

# 2. Split Text into Chunks
splitter = CharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_text(text)

# 3. Create Embeddings and Vector Store
embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")
db = FAISS.from_texts(chunks, embeddings)

# 4. Setup LLM and Prompt
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
prompt = ChatPromptTemplate.from_template(
    "Answer the question based only on the following context:\n"
    "{context}\n\nQuestion: {question}"
)

# 5. Retrieve and Generate
def ask_question(query):
    docs = db.similarity_search(query, k=3)
    context = "\n".join([d.page_content for d in docs])
    chain = prompt | llm
    response = chain.invoke({"context": context, "question": query})
    return response.content

print(ask_question("What are the main findings?"))

This code extracts text from a PDF, creates a local vector index using FAISS, and uses OpenAI’s GPT model to answer questions based on the retrieved context. Note that API keys must be set in environment variables for `ChatOpenAI` to work.

Common mistakes

  • Poor Chunking: Chunks too small lose context; too large dilute relevance. Use overlap to maintain continuity.
  • Ignoring Metadata: Failing to track which page or section a chunk came from makes citations impossible.
  • Static Retrieval: Using only simple similarity search can miss synonyms. Consider hybrid search (keyword + semantic).
  • Over-reliance on LLM: If the retrieval fails, the LLM will guess. Always verify if the context contains the answer before generating.

When to use it

ApproachBest ForLimitations
RAG (PDF Q&A) Large, static, or frequently updated document sets where accuracy and citation are critical. Complex setup; requires maintenance of vector databases.
Fine-Tuning Changing the model's style, tone, or domain-specific knowledge permanently. Expensive; does not easily update with new facts; high risk of hallucination if data is outdated.
Long Context Window Small documents that fit entirely within the LLM's token limit. Costly per query; performance degrades as context length increases ("lost in the middle").

Practice

Guided Exercise: Modify the example to print the source page number for each retrieved chunk. Hint: Store metadata during the `FAISS.from_texts` step or use `similarity_search_with_score`.

Challenge: Implement a "no answer" guardrail. If the similarity score of the top result is below a threshold, return "I cannot find this information in the document." instead of asking the LLM.

Quick check

Q: Why do we split PDFs into chunks before embedding?

A: Embedding models have token limits, and smaller chunks allow for more precise semantic matching between specific queries and relevant text segments.

Summary

PDF Q&A bots leverage RAG to combine the reasoning power of LLMs with the factual grounding of external documents. By structuring data into searchable vectors, you enable accurate, scalable, and cost-effective document interaction.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

PDF Q&A Bot with an LLM – FAQs

Quick answers about learning PDF Q&A Bot with an LLM in Data Science.

This free note from Coding Now Tech Institute explains PDF Q&A Bot with an LLM in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including PDF Q&A Bot with an LLM, is 100% free with no signup required.
With focused practice, most students grasp PDF Q&A Bot with an LLM in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now