Back to Data Science Notes
Topic #90

Retrieval-Augmented Generation (RAG)

By the end of this lesson, you will understand how Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) with external data to reduce hallucinations and improve accuracy.

What it is

Retrieval-Augmented Generation is an architecture that combines a generative model (like an LLM) with a retrieval system. Instead of relying solely on the model's pre-trained weights, RAG fetches relevant context from an external knowledge base (often a vector database) at query time. This retrieved information is then injected into the prompt, guiding the LLM to generate answers based on specific, up-to-date facts rather than general training data.

Mental Model: Think of RAG as giving the LLM an "open-book exam." The LLM knows how to read and synthesize information, but it needs the textbook (the vector database) to answer specific questions accurately.

Related terms: Embeddings, Vector Database, Chunking, Semantic Search, Prompt Engineering.

Why it matters

  • Reduces Hallucinations: By grounding responses in retrieved documents, the model is less likely to invent facts.
  • Access to Private Data: Allows LLMs to answer questions about internal company documents, codebases, or customer records without retraining.
  • Currency: Knowledge bases can be updated instantly, whereas retraining an LLM is slow and expensive.
  • Cost Efficiency: Smaller models can achieve high performance when provided with precise context, reducing inference costs.

Syntax or steps

A basic RAG pipeline involves three main stages:

  1. Ingestion: Split documents into chunks, convert them into numerical vectors (embeddings), and store them in a vector database.
  2. Retrieval: Convert the user's question into an embedding, search the vector database for the most similar chunks, and retrieve the top results.
  3. Generation: Construct a prompt containing the original question and the retrieved context, then send it to the LLM to generate the final answer.

Example

The following Python example uses langchain and chromadb to demonstrate a minimal RAG workflow. Note that API keys are required for the LLM and embedding model.

from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import TextLoader

# 1. Load and split documents
loader = TextLoader("state_of_the_union.txt")
documents = loader.load()
text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
splits = text_splitter.split_documents(documents)

# 2. Create embeddings and vector store
embedding_function = OpenAIEmbeddings()
vectorstore = Chroma.from_documents(documents=splits, embedding=embedding_function)

# 3. Define retriever
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})

# 4. Define LLM and Prompt
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
template = """Answer the question based only on the following context:
{context}

Question: {question}
"""
prompt = ChatPromptTemplate.from_template(template)

def format_docs(docs):
    return "\n\n".join([d.page_content for d in docs])

# 5. Chain components together
rag_chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | prompt
    | llm
)

# 6. Query
response = rag_chain.invoke("What did the president say about technology?")
print(response.content)

Explanation: The code first loads text and splits it into manageable chunks. It then creates a vector store using OpenAI embeddings. A retriever is configured to find the top 4 relevant chunks. Finally, a chain is built where the retrieved context and the user's question are formatted into a prompt template before being sent to the LLM.

Common mistakes

  • Poor Chunking Strategy: Chunks that are too large dilute relevance; chunks that are too small lose context. Use overlap to maintain semantic continuity.
  • Ignoring Metadata: Failing to store source metadata makes it impossible to cite references or filter searches by date or author.
  • Over-reliance on Top-K: Retrieving too many irrelevant chunks can confuse the LLM ("lost in the middle" phenomenon). Tune k carefully.
  • Static Prompts: Not instructing the LLM to admit if the answer isn't in the context leads to hallucinations even with RAG.

When to use it

Approach Best For Limitations
RAG Dynamic data, private knowledge, factual Q&A, citations. Dependent on retrieval quality; latency increases with search step.
Fine-Tuning Changing style, tone, or teaching new skills/formats. Expensive; knowledge becomes stale quickly; hard to update.

Practice

Guided Exercise: Modify the example above to change the chunk_size to 500 and observe how the retrieved context changes for a specific query.

Challenge: Implement a simple evaluation metric by checking if the generated answer contains keywords found in the retrieved documents. Hint: Use string matching or cosine similarity between the answer and the context.

Quick check

Q: Why is RAG preferred over fine-tuning for answering questions about recent news?

A: Because RAG allows instant updates to the knowledge base (vector DB) without the time and cost of retraining the model, ensuring the LLM has access to the latest information.

Summary

RAG bridges the gap between static LLM weights and dynamic external knowledge by retrieving relevant context at runtime. It is essential for building accurate, up-to-date, and verifiable AI applications that rely on specific datasets.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Retrieval-Augmented Generation (RAG) – FAQs

Quick answers about learning Retrieval-Augmented Generation (RAG) in Data Science.

This free note from Coding Now Tech Institute explains Retrieval-Augmented Generation (RAG) in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Retrieval-Augmented Generation (RAG), is 100% free with no signup required.
With focused practice, most students grasp Retrieval-Augmented Generation (RAG) in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now