EXTENDING THE CHATBOT
Hello, students! It is wonderful to connect with you again. In our first adventure, we embarked on an exciting journey to construct a basic AI chatbot using Python, Streamlit, LangChain, and Hugging Face models. We saw how our chatbot could hold conversations and even search the internet for information. That was a fantastic start, and you all did a brilliant job mastering those foundational concepts.
Now, imagine this: Your chatbot is smart, but its knowledge is mostly what it learned during its training, or what it can find on the internet in real-time. What if you have a special collection of documents, like your school textbooks, research papers, or even notes from your favorite classes, that you want your chatbot to be an expert on? How can we teach our chatbot to deeply understand and answer questions specifically from your documents, without having to retrain a massive AI model every time?
This is exactly what we are going to tackle in Part 2 of our tutorial! We will learn about an awesome technique called Retrieval Augmented Generation, or RAG for short. RAG is like giving your chatbot a personal library and teaching it how to quickly find the right book (or even just the right page) to answer any question you throw at it. This way, your chatbot becomes not only a great conversationalist but also a knowledgeable expert on your chosen topics.
By the end of this part, your chatbot will be able to answer questions based on its own internal knowledge, search the web, and consult your custom document library. How exciting is that? Let us dive in!
The Challenge: Limited Knowledge and "Hallucinations"
Remember how we talked about Large Language Models (LLMs) learning from vast amounts of text data? While this makes them incredibly versatile, it also comes with a few limitations.
First, an LLM's knowledge is frozen at the time it was trained. It does not automatically know about events that happened yesterday or new discoveries made last week. This is why we added the internet search feature in Part 1, to give it access to current information.
Second, sometimes LLMs can "hallucinate." This means they might confidently make up facts or provide incorrect information, especially when asked about very specific or obscure topics that were not well-represented in their training data. It is not because they are trying to trick us, but because they are trying their best to generate a coherent and plausible response based on patterns they have learned, even if they do not have the exact factual basis.
To make our chatbot truly reliable and an expert on specific subjects, we need a way to provide it with accurate, up-to-date, and domain-specific information before it generates a response. This is where Retrieval Augmented Generation comes to the rescue!
Introducing Retrieval Augmented Generation (RAG)
Think of RAG as a two-step process for your chatbot:
- Retrieval: When you ask a question, the chatbot first acts like a super-fast librarian. It quickly scans your personal document library to find all the pieces of information that are most relevant to your question.
- Generation: Once it has found those relevant pieces, it then uses its powerful language generation abilities to formulate an answer, but now it has the precise context from your documents right in front of it. It is like having the answer key before taking a test!
This process makes your chatbot more accurate, reduces hallucinations, and allows it to specialize in any knowledge base you provide.
To implement RAG, we will need to learn about a few new concepts and tools, but do not worry, we will go through each one step-by-step.
Setting Up Your RAG Toolkit
Before we start coding, we need to ensure our Python environment has all the necessary libraries. If you have not already, please install these using pip:
pip install langchain langchain-community sentence-transformers faiss-cpu pypdf
langchainandlangchain-communityare our trusty friends for orchestrating LLMs and various components.sentence-transformershelps us turn text into numerical representations.faiss-cpuis a library that helps us store and quickly search these numerical representations.pypdfhelps us read PDF documents, which are a common format for knowledge bases.
Step 1: Gathering Your Knowledge Base and Breaking Down Big Ideas (Chunking)
Our chatbot needs documents to learn from. For this tutorial, let us imagine you have a PDF document named my_science_notes.pdf filled with all your amazing science notes. You can replace this with any PDF document you wish to use as your chatbot's knowledge base.
Now, imagine giving a whole textbook to someone and asking them to find one specific answer. It would take a long time! LLMs face a similar problem. They have a limit to how much text they can process at once. So, instead of giving the entire document to the LLM, we break our large documents into smaller, more manageable pieces called "chunks." This process is known as chunking.
Why chunking?
- Manageable Size: Smaller chunks fit within the LLM's input limit.
- Relevance: When searching for information, it is easier to find a few relevant small chunks than to sift through a massive document.
- Efficiency: Processing smaller pieces of text is faster and uses less computational power.
We will use LangChain's RecursiveCharacterTextSplitter to intelligently split our documents. This splitter tries to keep related sentences together by splitting on different characters (like newlines, then spaces, then commas) in a hierarchical way.
Here is how you can load your document and split it into chunks:
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
# First, specify the path to your document.
# Make sure 'my_science_notes.pdf' is in the same directory as your Python script,
# or provide the full path to the file.
document_path = "my_science_notes.pdf"
# Initialize the PDF loader. This tool knows how to read PDF files.
print(f"Loading document from: {document_path}")
loader = PyPDFLoader(document_path)
# Load the entire document. The loader reads the PDF and extracts its text content.
# It returns a list of 'Document' objects, where each object might represent a page.
documents = loader.load()
print(f"Successfully loaded {len(documents)} pages from the document.")
# Initialize the text splitter. This is like setting up rules for how to break down the text.
# 'chunk_size' determines the maximum number of characters in each piece.
# 'chunk_overlap' specifies how many characters each chunk will share with the previous one.
# Overlapping helps ensure that context is not lost when a key piece of information
# falls exactly on a chunk boundary.
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1000, # Each chunk will aim to be around 1000 characters long
chunk_overlap=200, # Chunks will overlap by 200 characters to maintain context
length_function=len, # We are measuring chunk size by character count
)
# Split the loaded documents into smaller, manageable chunks.
print("Splitting documents into smaller chunks...")
chunks = text_splitter.split_documents(documents)
print(f"Original document split into {len(chunks)} chunks.")
# Let us print the first chunk to see what it looks like.
# This helps us understand how the text has been divided.
if chunks:
print("\n--- Example of a document chunk ---")
print(chunks[0].page_content)
print("----------------------------------")
else:
print("No chunks were created. Please check your document and splitter settings.")
Step 2: Turning Words into Numbers (Embeddings)
Now that we have our document broken into chunks, how does our chatbot find the most relevant chunks for a given question? It cannot just read every chunk and compare it word-for-word. That would be too slow!
This is where embeddings come in. An embedding is a numerical representation of a piece of text. Imagine taking a sentence, a paragraph, or even a whole chunk, and turning it into a long list of numbers (a vector). The magic is that texts with similar meanings will have similar lists of numbers. In a high-dimensional space, these similar numerical lists will be "close" to each other.
So, when you ask a question, we will turn your question into an embedding (a list of numbers). Then, we will compare your question's embedding to the embeddings of all our document chunks. The chunks whose embeddings are numerically "closest" to your question's embedding are the most relevant ones!
We will use a pre-trained embedding model from Hugging Face, specifically sentence-transformers/all-MiniLM-L6-v2, which is excellent for this task and relatively lightweight.
Here is how you generate embeddings for your chunks:
from langchain_community.embeddings import HuggingFaceEmbeddings
# Specify the name of the pre-trained embedding model we want to use.
# This model is designed to convert sentences and paragraphs into meaningful numerical vectors.
embedding_model_name = "sentence-transformers/all-MiniLM-L6-v2"
# Initialize the embedding model.
# The first time you run this, the model will be downloaded to your computer.
# This might take a moment, depending on your internet connection.
print(f"\nInitializing embedding model: {embedding_model_name}")
embeddings = HuggingFaceEmbeddings(model_name=embedding_model_name)
print("Embedding model initialized successfully.")
# You can test the embedding model by converting a simple sentence into a vector.
# This helps confirm that the model is working as expected.
sample_text = "The quick brown fox jumps over the lazy dog."
sample_vector = embeddings.embed_query(sample_text)
print(f"Example: Embedding vector for '{sample_text}' has length: {len(sample_vector)}")
# The length of the vector indicates the dimensionality of the embedding space.
Step 3: Storing and Finding Knowledge (Vector Database)
We have our chunks, and we know how to turn them into numerical embeddings. Now, we need a place to store all these embeddings efficiently and a way to quickly search through them to find the closest ones to our question's embedding. This special storage and search system is called a vector database (or vector store).
For our tutorial, we will use FAISS (Facebook AI Similarity Search). FAISS is not a full-fledged database like SQL or NoSQL, but it is an incredibly fast and efficient library specifically designed for similarity search on large collections of vectors. It is perfect for our needs because it is easy to set up locally and works entirely in memory (or can be saved to disk).
Here is how you create a FAISS vector store from your chunks and embeddings, and how you can save it for future use:
from langchain_community.vectorstores import FAISS
# Ensure 'chunks' and 'embeddings' from the previous steps are available.
# If you are running this code separately, you would need to re-run the chunking
# and embedding model initialization steps first.
# Create a FAISS vector store.
# This step takes all your document chunks, converts each one into an embedding
# using the 'embeddings' model we initialized, and then builds an efficient
# index within FAISS for fast similarity searching.
print("\nCreating FAISS vector store from document chunks...")
vector_store = FAISS.from_documents(chunks, embeddings)
print("FAISS vector store created successfully.")
# It is a good practice to save your vector store to disk.
# This way, you do not have to re-process all your documents and generate embeddings
# every single time you start your chatbot. You can just load the saved index.
vector_store_path = "faiss_index_chatbot"
print(f"Saving vector store to local disk at: {vector_store_path}")
vector_store.save_local(vector_store_path)
print("Vector store saved successfully. You can now load it directly next time.")
# To demonstrate how to load it back, here is the code you would use
# in a new session or after restarting your application:
print(f"\nDemonstrating how to load the vector store from: {vector_store_path}")
# When loading, you must provide the embedding model again, as FAISS needs it
# to understand how to interpret and compare vectors.
loaded_vector_store = FAISS.load_local(
vector_store_path,
embeddings,
# This parameter is important for security; it allows loading objects
# that might contain custom Python code. For our local use, it is fine,
# but be cautious with untrusted sources.
allow_dangerous_deserialization=True
)
print("Vector store loaded successfully for demonstration.")
# You can test the loaded vector store by performing a simple similarity search.
# This will find chunks that are most similar to your query.
test_query = "What are the main components of a cell?"
print(f"\nSearching for documents related to: '{test_query}'")
retrieved_docs = loaded_vector_store.similarity_search(test_query, k=2) # Retrieve top 2 similar documents
print("Retrieved documents (showing content of the first one):")
if retrieved_docs:
print(retrieved_docs[0].page_content)
else:
print("No documents retrieved for the test query.")
Step 4: Bringing it All Together: RAG in Action with Your Chatbot
Now for the exciting part: integrating our RAG system into the chatbot we built in Part 1! Our goal is to make sure that when a user asks a question, our chatbot first retrieves relevant information from our FAISS vector store and then uses that information to formulate a more informed answer.
Since our Part 1 chatbot already had conversational memory, we will use LangChain's ConversationalRetrievalChain. This chain is specifically designed to handle both retrieval and chat history, making our chatbot even smarter.
Here is how you would modify your existing chatbot code (from Part 1) to incorporate RAG:
from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain_community.llms import HuggingFacePipeline
from transformers import pipeline
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
import streamlit as st # Assuming your chatbot UI is built with Streamlit
# --- 1. Load the Language Model (LLM) from Part 1 ---
# This part should be similar to how you set up your LLM in Part 1.
# We are using a Hugging Face model (like GPT-2) via a pipeline.
print("\nInitializing the Large Language Model (LLM) pipeline...")
llm_pipeline = pipeline(
"text-generation",
model="gpt2", # Replace with the specific model you used in Part 1 if different
tokenizer="gpt2",
max_new_tokens=500, # Maximum number of tokens the LLM will generate in its response
device=-1, # -1 for CPU, 0 for GPU if you have one and PyTorch is configured for it
)
llm = HuggingFacePipeline(pipeline=llm_pipeline)
print("LLM pipeline initialized.")
# --- 2. Load the Embedding Model ---
# We need the same embedding model to load our vector store and to embed new queries.
print("Initializing embedding model for loading vector store...")
embedding_model_name = "sentence-transformers/all-MiniLM-L6-v2"
embeddings = HuggingFaceEmbeddings(model_name=embedding_model_name)
print("Embedding model initialized.")
# --- 3. Load Your FAISS Vector Store ---
# This loads the knowledge base you created and saved in the previous step.
vector_store_path = "faiss_index_chatbot"
print(f"Loading FAISS vector store from: {vector_store_path}")
try:
loaded_vector_store = FAISS.load_local(
vector_store_path,
embeddings,
allow_dangerous_deserialization=True
)
print("Vector store loaded successfully.")
except Exception as e:
print(f"Error loading vector store: {e}")
print("Please ensure you have run Step 3 to create and save the 'faiss_index_chatbot'.")
# Handle the error, perhaps by exiting or using a fallback mechanism
st.error("Could not load knowledge base. Please ensure it is created.")
st.stop() # Stop Streamlit execution if the knowledge base isn't available
# --- 4. Create a Retriever from the Vector Store ---
# The retriever is the component that knows how to search your vector store
# for relevant documents based on a query.
# 'k=4' means it will retrieve the top 4 most relevant document chunks.
retriever = loaded_vector_store.as_retriever(search_kwargs={"k": 4})
print("Retriever created from the vector store.")
# --- 5. Initialize Conversational Memory (from Part 1) ---
# This memory object will keep track of the conversation history,
# allowing the chatbot to remember previous turns.
print("Initializing conversational memory...")
memory = ConversationBufferMemory(
memory_key="chat_history", # This key tells LangChain where to find the chat history
return_messages=True # We want the memory to return actual message objects
)
print("Conversational memory initialized.")
# --- 6. Create the Conversational RAG Chain ---
# This is the core of our RAG chatbot. It combines the LLM, the retriever,
# and the conversational memory into a single, powerful chain.
# 'condense_question_llm=llm' means the LLM itself will be used to
# rephrase the user's question, taking into account the chat history,
# to make it a standalone question for better retrieval.
# 'return_source_documents=True' is very useful for debugging and for
# showing the user where the information came from.
print("Creating the Conversational Retrieval Chain...")
qa_chain = ConversationalRetrievalChain.from_llm(
llm=llm,
retriever=retriever,
memory=memory,
condense_question_llm=llm, # Use the same LLM to condense questions
return_source_documents=True
)
print("Conversational Retrieval Chain created. Your RAG chatbot is ready!")
# --- 7. Integrate into your Streamlit App (Conceptual Example) ---
# This section shows how you would typically integrate 'qa_chain' into your
# existing Streamlit application's chat loop.
# Assuming your Streamlit app has a way to get user input and display responses:
# if "messages" not in st.session_state:
# st.session_state.messages = []
# for message in st.session_state.messages:
# with st.chat_message(message["role"]):
# st.markdown(message["content"])
# user_query = st.chat_input("Ask your RAG-powered chatbot:")
# if user_query:
# st.session_state.messages.append({"role": "user", "content": user_query})
# with st.chat_message("user"):
# st.markdown(user_query)
# with st.chat_message("assistant"):
# # Instead of calling your old conversation chain, call the new qa_chain
# # The input key for ConversationalRetrievalChain is "question"
# response = qa_chain.invoke({"question": user_query})
# bot_answer = response["answer"]
# st.markdown(bot_answer)
# # Optionally, display the source documents
# if response["source_documents"]:
# st.subheader("Sources:")
# for i, doc in enumerate(response["source_documents"]):
# st.text(f"Source {i+1}: {doc.metadata.get('source', 'Unknown source')}")
# st.text(doc.page_content[:200] + "...") # Show first 200 chars of source
# st.session_state.messages.append({"role": "assistant", "content": bot_answer})
print("\nConceptual Streamlit integration shown. You will replace your old chatbot call with 'qa_chain.invoke'.")
print("Remember to adapt your Streamlit UI to display source documents if desired.")
In the code above, we first load our LLM and the embedding model, just like before. Then, we load our pre-built FAISS vector store. We turn this vector store into a retriever, which is LangChain's way of saying "the thing that fetches documents." We also bring back our ConversationBufferMemory to keep our chatbot conversational.
Finally, we create the ConversationalRetrievalChain. When you ask a question, this chain will:
- Look at your current question and the chat history to form a clear, standalone question.
- Use this standalone question to query the
retrieverand find the most relevant chunks from yourfaiss_index_chatbot. - Combine these retrieved chunks with the chat history and your original question, and send all of this context to the LLM.
- The LLM then generates an answer based on all this rich information.
The return_source_documents=True parameter is incredibly useful because it makes the chain return not just the answer but also the specific document chunks it used to generate that answer. This helps you (and your users) verify the information and understand its origin.
Step 5: Testing and Refining Your RAG Chatbot
Now that your chatbot is equipped with RAG, it is time to test it out!
Run your Streamlit application and try asking questions that are specifically related to the content of your my_science_notes.pdf document.
For example, if your science notes are about photosynthesis, ask: "What are the main stages of photosynthesis?" or "What role does chlorophyll play?"
Compare the answers you get now with the answers you might have received from the chatbot in Part 1 (without RAG). You should notice that the RAG-powered chatbot provides more precise, detailed, and accurate answers directly from your document.
You might also want to experiment with:
chunk_sizeandchunk_overlap: How do different chunking strategies affect the quality of retrieved documents and answers?kinretriever.as_retriever(search_kwargs={"k": 4}): What happens if you retrieve more or fewer documents? Does retrieving too many confuse the LLM, or does retrieving too few miss important context?- The content of your documents: The quality of your RAG system heavily depends on the quality and comprehensiveness of your knowledge base.
Remember, building AI applications is an iterative process. You will continuously learn, experiment, and refine your creation.
Conclusion: Your Smart, Knowledgeable Chatbot!
Congratulations! You have successfully extended your basic AI chatbot into a powerful, knowledge-augmented system using Retrieval Augmented Generation. You have learned how to:
- Break down large documents into manageable chunks.
- Convert text into numerical embeddings that capture meaning.
- Store and efficiently search these embeddings using a vector database like FAISS.
- Integrate this retrieval capability with your LLM and conversational memory using LangChain.
Your chatbot is now not just a general conversationalist but also a specialized expert, capable of drawing information directly from your custom knowledge base. This opens up a world of possibilities for creating AI assistants tailored to specific subjects, industries, or even just your personal learning needs.
Keep experimenting, keep learning, and keep building! The world of AI is vast and exciting, and you are now equipped with even more powerful tools to explore it. What will you teach your chatbot next? The possibilities are truly endless!
The Complete RAG Chatbot Application Code
Here is the full Python script for your RAG-powered Streamlit chatbot. This code combines all the steps we covered: loading documents, chunking, creating embeddings, building a FAISS vector store, and integrating it with a conversational LLM and memory.
import streamlit as st
import os
import torch
from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM
from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain_community.llms import HuggingFacePipeline
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
# --- Configuration ---
# Define the path for your knowledge base PDF file.
# IMPORTANT: Place your 'my_science_notes.pdf' file in the same directory as this script.
# If the file doesn't exist, the script will prompt you to create it.
KNOWLEDGE_BASE_PDF_PATH = "my_science_notes.pdf"
# Define the path where your FAISS vector store will be saved/loaded.
FAISS_INDEX_PATH = "faiss_index_chatbot"
# Define the Hugging Face model for embeddings.
EMBEDDING_MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
# Define the Hugging Face model for the Large Language Model (LLM).
# This should be the same model you used in Part 1.
# Examples: "gpt2", "distilgpt2", "microsoft/DialoGPT-small"
LLM_MODEL_NAME = "gpt2"
# --- Helper Function to Initialize Knowledge Base ---
def initialize_knowledge_base():
"""
Loads documents, chunks them, creates embeddings, and builds/saves a FAISS vector store.
If the FAISS index already exists, it loads it.
"""
# Initialize embeddings model first, as it's needed for both creating and loading the FAISS index.
print(f"Initializing embedding model: {EMBEDDING_MODEL_NAME}")
embeddings = HuggingFaceEmbeddings(model_name=EMBEDDING_MODEL_NAME)
print("Embedding model initialized.")
if os.path.exists(FAISS_INDEX_PATH):
print(f"Loading existing FAISS vector store from: {FAISS_INDEX_PATH}")
try:
# When loading, we must provide the embedding model again, as FAISS needs it
# to understand how to interpret and compare vectors.
vector_store = FAISS.load_local(FAISS_INDEX_PATH, embeddings, allow_dangerous_deserialization=True)
print("Vector store loaded successfully.")
return vector_store
except Exception as e:
st.error(f"Error loading existing vector store: {e}. Attempting to rebuild.")
# If loading fails, proceed to rebuild
pass
print("FAISS vector store not found or failed to load. Initializing from documents...")
if not os.path.exists(KNOWLEDGE_BASE_PDF_PATH):
st.error(f"Error: Knowledge base PDF '{KNOWLEDGE_BASE_PDF_PATH}' not found.")
st.info("Please create a PDF file named 'my_science_notes.pdf' in the same directory as this script, "
"filled with information you want your chatbot to learn. For example, you could put notes about "
"biology, chemistry, or physics in it.")
st.stop() # Stop the Streamlit app if the PDF is missing
print(f"Loading document from: {KNOWLEDGE_BASE_PDF_PATH}")
# Initialize the PDF loader. This tool knows how to read PDF files.
loader = PyPDFLoader(KNOWLEDGE_BASE_PDF_PATH)
# Load the entire document. The loader reads the PDF and extracts its text content.
# It returns a list of 'Document' objects, where each object might represent a page.
documents = loader.load()
print(f"Successfully loaded {len(documents)} pages from the document.")
# Initialize the text splitter. This is like setting up rules for how to break down the text.
# 'chunk_size' determines the maximum number of characters in each piece.
# 'chunk_overlap' specifies how many characters each chunk will share with the previous one.
# Overlapping helps ensure that context is not lost when a key piece of information
# falls exactly on a chunk boundary.
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1000, # Each chunk will aim to be around 1000 characters long
chunk_overlap=200, # Chunks will overlap by 200 characters to maintain context
length_function=len, # We are measuring chunk size by character count
)
# Split the loaded documents into smaller, manageable chunks.
chunks = text_splitter.split_documents(documents)
print(f"Original document split into {len(chunks)} chunks.")
print("Creating FAISS vector store from document chunks...")
# This step takes all your document chunks, converts each one into an embedding
# using the 'embeddings' model we initialized, and then builds an efficient
# index within FAISS for fast similarity searching.
vector_store = FAISS.from_documents(chunks, embeddings)
print("FAISS vector store created successfully.")
print(f"Saving vector store to local disk at: {FAISS_INDEX_PATH}")
# It is a good practice to save your vector store to disk.
# This way, you do not have to re-process all your documents and generate embeddings
# every single time you start your chatbot. You can just load the saved index.
vector_store.save_local(FAISS_INDEX_PATH)
print("Vector store saved successfully.")
return vector_store
# --- Streamlit UI Setup ---
# Configure the Streamlit page with a title, icon, and layout.
st.set_page_config(page_title="RAG-Powered Chatbot", page_icon="📚", layout="wide")
st.title("📚 Your RAG-Powered Chatbot (Part 2)")
st.write("Hello, future AI developers! This chatbot now uses Retrieval Augmented Generation (RAG) to answer questions based on your custom knowledge base!")
st.markdown("---")
# --- Initialize Components (using st.session_state to avoid re-running on every interaction) ---
# Streamlit's `st.session_state` is crucial here. It allows us to store objects
# (like our LLM, vector store, and chains) so they are not re-initialized
# every time the user interacts with the app (e.g., typing a new message).
# Initialize LLM (Large Language Model)
if "llm" not in st.session_state:
print(f"\nInitializing LLM pipeline: {LLM_MODEL_NAME}...")
try:
# Load the tokenizer and model for the specified LLM.
tokenizer = AutoTokenizer.from_pretrained(LLM_MODEL_NAME)
model = AutoModelForCausalLM.from_pretrained(LLM_MODEL_NAME)
# Determine the device to run the LLM on (GPU if available, otherwise CPU).
device = 0 if torch.cuda.is_available() else -1
print(f"Using device: {'cuda' if device == 0 else 'cpu'}")
# Create a Hugging Face pipeline for text generation.
pipe = pipeline(
"text-generation",
model=model,
tokenizer=tokenizer,
max_new_tokens=500, # Maximum number of tokens the LLM will generate in its response
device=device, # Assign the model to the detected device
do_sample=True, # Enable sampling for more varied responses
temperature=0.7, # Control creativity (lower for more deterministic, higher for more creative)
top_k=50, # Limit sampling to top_k most likely words
top_p=0.95 # Limit sampling to words that make up top_p probability mass
)
# Wrap the Hugging Face pipeline in LangChain's HuggingFacePipeline class.
st.session_state.llm = HuggingFacePipeline(pipeline=pipe)
print("LLM pipeline initialized.")
except Exception as e:
st.error(f"Failed to load LLM ({LLM_MODEL_NAME}). Please ensure it's installed and accessible. "
f"You might need to run `pip install transformers torch` and check your internet connection. Error: {e}")
st.stop() # Stop the app if the LLM cannot be loaded.
# Initialize Vector Store (Knowledge Base)
if "vector_store" not in st.session_state:
st.session_state.vector_store = initialize_knowledge_base()
# Create Retriever
if "retriever" not in st.session_state:
# The retriever is the component that knows how to search your vector store
# for relevant documents based on a query.
# 'k=4' means it will retrieve the top 4 most relevant document chunks.
st.session_state.retriever = st.session_state.vector_store.as_retriever(search_kwargs={"k": 4})
print("Retriever created from the vector store.")
# Initialize Conversational Memory
if "memory" not in st.session_state:
# This memory object will keep track of the conversation history,
# allowing the chatbot to remember previous turns.
st.session_state.memory = ConversationBufferMemory(
memory_key="chat_history", # This key tells LangChain where to find the chat history
return_messages=True # We want the memory to return actual message objects
)
print("Conversational memory initialized.")
# Create Conversational Retrieval Chain
if "qa_chain" not in st.session_state:
# This is the core of our RAG chatbot. It combines the LLM, the retriever,
# and the conversational memory into a single, powerful chain.
# 'condense_question_llm=st.session_state.llm' means the LLM itself will be used to
# rephrase the user's question, taking into account the chat history,
# to make it a standalone question for better retrieval.
# 'return_source_documents=True' is very useful for debugging and for
# showing the user where the information came from.
print("Creating the Conversational Retrieval Chain...")
st.session_state.qa_chain = ConversationalRetrievalChain.from_llm(
llm=st.session_state.llm,
retriever=st.session_state.retriever,
memory=st.session_state.memory,
condense_question_llm=st.session_state.llm, # Use the same LLM to condense questions
return_source_documents=True
)
print("Conversational Retrieval Chain created. Your RAG chatbot is ready!")
# --- Display Chat History ---
# If there are no messages in the session state, initialize an empty list.
if "messages" not in st.session_state:
st.session_state.messages = []
# Iterate through all stored messages and display them in the chat interface.
for message in st.session_state.messages:
with st.chat_message(message["role"]):
st.markdown(message["content"])
# If a message has associated sources, display them in an expandable section.
if "sources" in message and message["sources"]:
with st.expander("Sources Used"):
for i, source in enumerate(message["sources"]):
st.text(f"Source {i+1}: {source}")
# --- Handle User Input ---
# Create a text input box for the user to type their questions.
user_query = st.chat_input("Ask your RAG-powered chatbot a question about your knowledge base!")
# If the user has entered a query:
if user_query:
# Add the user's query to the chat history.
st.session_state.messages.append({"role": "user", "content": user_query})
with st.chat_message("user"):
st.markdown(user_query)
# Display a "Thinking..." spinner while the chatbot processes the query.
with st.chat_message("assistant"):
with st.spinner("Thinking..."):
try:
# Call the RAG-powered conversational chain to get a response.
response = st.session_state.qa_chain.invoke({"question": user_query})
bot_answer = response["answer"]
source_documents = response.get("source_documents", [])
# Display the chatbot's answer.
st.markdown(bot_answer)
# Display source documents if available.
if source_documents:
with st.expander("Sources Used"):
for i, doc in enumerate(source_documents):
# Extract filename from the full path for cleaner display.
file_name = os.path.basename(doc.metadata.get('source', 'Unknown file'))
st.markdown(f"**Source {i+1}** (Page: {doc.metadata.get('page', 'N/A')}, "
f"File: {file_name}):")
st.text(doc.page_content[:300] + "...") # Show first 300 characters of source content
# Store the assistant's message and sources in session state for display in future turns.
st.session_state.messages.append({"role": "assistant", "content": bot_answer,
"sources": [f"Page {doc.metadata.get('page', 'N/A')} from "
f"{os.path.basename(doc.metadata.get('source', 'Unknown file'))}"
for doc in source_documents]})
except Exception as e:
# Handle any errors that occur during the RAG process.
st.error(f"An error occurred while processing your request: {e}")
st.session_state.messages.append({"role": "assistant", "content": f"I'm sorry, I encountered an error: {e}"})
How to Build, Deploy, and Start Your RAG Chatbot Application
Now that you have the complete code, let's get it up and running! This process involves a few straightforward steps:
Step 1: Prepare Your Environment (Install Dependencies)
First, you need to make sure your Python environment has all the necessary libraries installed. Open your terminal or command prompt and run the following command:
pip install langchain langchain-community sentence-transformers faiss-cpu pypdf transformers torch streamlit
langchainandlangchain-community: These are the core libraries for building LLM applications, helping us orchestrate all the components.sentence-transformers: This library provides the embedding model we use to convert text into numerical vectors.faiss-cpu: This is Facebook AI Similarity Search, a super-fast library for storing and searching our text embeddings.pypdf: This library allows our application to read and extract text from PDF documents.transformersandtorch: These are essential for loading and running the Hugging Face Large Language Model (LLM) you chose (like GPT-2).torchis the underlying deep learning framework.streamlit: This is the framework we use to create the interactive web-based user interface for our chatbot.
Step 2: Save the Application Code
- Create a new file: Open a plain text editor (like VS Code, Sublime Text, Notepad++, or even a simple text editor).
- Copy the code: Copy the entire Python script provided above.
- Save the file: Save the copied code into a file named
rag_chatbot_app.py(or any other.pyname you prefer) in a folder on your computer. Remember the location of this folder!
Step 3: Create Your Knowledge Base Document
Your chatbot needs something to learn from!
- Create a PDF file: In the same folder where you saved
rag_chatbot_app.py, create a PDF document namedmy_science_notes.pdf.Content: Fill this PDF with information you want your chatbot to be an expert on. For example, you could:
- Copy-paste notes from your science textbook about topics like photosynthesis, the solar system, or human anatomy.
- Write a short story or a fictional history.
- Include details about a specific project or hobby.
How to create a PDF: You can use a word processor (like Microsoft Word, Google Docs, LibreOffice Writer) to type out your content and then use its "Save As PDF" or "Print to PDF" function. Make sure the saved file is named
my_science_notes.pdf.Example Content for
my_science_notes.pdf(you can copy this into a document and save as PDF):The process of photosynthesis is how plants convert light energy into chemical energy, which is stored in glucose. This process primarily occurs in the chloroplasts, specifically using chlorophyll, the green pigment. Photosynthesis requires carbon dioxide, water, and sunlight. It produces glucose (sugar) and oxygen. The main stages are the light-dependent reactions and the light-independent reactions (Calvin Cycle). During light-dependent reactions, light energy is captured by chlorophyll and converted into ATP and NADPH, releasing oxygen as a byproduct. The Calvin Cycle then uses this ATP and NADPH to convert carbon dioxide into glucose. The human heart is a muscular organ that pumps blood throughout the body. It has four chambers: two atria and two ventricles. The right side of the heart pumps deoxygenated blood to the lungs, while the left side pumps oxygenated blood to the rest of the body. The average adult heart beats about 60 to 100 times per minute.
Step 4: Run Your Streamlit Application
Now for the moment of truth!
Open your terminal/command prompt: Navigate to the folder where you saved
rag_chatbot_app.pyandmy_science_notes.pdf. You can do this using thecdcommand (e.g.,cd path/to/your/folder).Start the Streamlit app: Once you are in the correct directory, run the following command:
streamlit run rag_chatbot_app.py- First Run: The first time you run this, it might take a few minutes. Streamlit will start, and the application will:
- Download the Hugging Face
gpt2model (if not already cached). - Download the
sentence-transformers/all-MiniLM-L6-v2embedding model. - Load your
my_science_notes.pdf, chunk it, create embeddings, and build the FAISS vector store. This process will print messages to your terminal indicating its progress. - Finally, it will save the FAISS index to a folder named
faiss_index_chatbotin your application directory.
- Download the Hugging Face
- First Run: The first time you run this, it might take a few minutes. Streamlit will start, and the application will:
Access the application: Once it finishes loading, your web browser should automatically open a new tab displaying your RAG-powered chatbot. If it does not, look for a URL in your terminal output that starts with
http://localhost:8501and open it manually.
Step 5: Interact with Your Chatbot!
You now have a fully functional RAG chatbot!
- Ask questions from your PDF: Type questions into the chat input box that are specifically related to the content you put in
my_science_notes.pdf.- Try questions like: "What is photosynthesis and where does it occur?" or "What are the main stages of photosynthesis?"
- Or, if you included information about the heart: "How many chambers does the human heart have?" or "What is the function of the right side of the heart?"
- Observe the responses: You should see your chatbot provide answers that are directly informed by the content of your PDF.
- Check the sources: Below the chatbot's answer, you will find an expandable section labeled "Sources Used." Click on it to see which specific chunks from your PDF the chatbot used to formulate its response. This is a powerful feature for verifying information!
- Test general knowledge: You can also ask general questions not found in your PDF. The LLM will still try to answer these based on its pre-trained knowledge.
What Happens Next Time You Run the App?
Because we added the logic to save and load the FAISS index, when you run streamlit run rag_chatbot_app.pyagain, it will detect the faiss_index_chatbot folder, load the pre-built vector store, and skip the time-consuming process of re-chunking and re-embedding your PDF. This makes subsequent startups much faster!
Congratulations! You have successfully built and deployed a sophisticated RAG-powered chatbot. This is a significant step in your AI journey, giving you the power to create intelligent agents that are experts in specific domains. Keep experimenting, keep learning, and enjoy your smarter chatbot!
No comments:
Post a Comment