Sunday, October 11, 2026

ADDING RAG TO YOUR LLM CHATBOT - PART 2/2



EXTENDING THE CHATBOT

Hello, students! It is wonderful to connect with you again. In our first adventure, we embarked on an exciting journey to construct a basic AI chatbot using Python, Streamlit, LangChain, and Hugging Face models. We saw how our chatbot could hold conversations and even search the internet for information. That was a fantastic start, and you all did a brilliant job mastering those foundational concepts.

Now, imagine this: Your chatbot is smart, but its knowledge is mostly what it learned during its training, or what it can find on the internet in real-time. What if you have a special collection of documents, like your school textbooks, research papers, or even notes from your favorite classes, that you want your chatbot to be an expert on? How can we teach our chatbot to deeply understand and answer questions specifically from your documents, without having to retrain a massive AI model every time?

This is exactly what we are going to tackle in Part 2 of our tutorial! We will learn about an awesome technique called Retrieval Augmented Generation, or RAG for short. RAG is like giving your chatbot a personal library and teaching it how to quickly find the right book (or even just the right page) to answer any question you throw at it. This way, your chatbot becomes not only a great conversationalist but also a knowledgeable expert on your chosen topics.

By the end of this part, your chatbot will be able to answer questions based on its own internal knowledge, search the web, and consult your custom document library. How exciting is that? Let us dive in!


The Challenge: Limited Knowledge and "Hallucinations"

Remember how we talked about Large Language Models (LLMs) learning from vast amounts of text data? While this makes them incredibly versatile, it also comes with a few limitations.

First, an LLM's knowledge is frozen at the time it was trained. It does not automatically know about events that happened yesterday or new discoveries made last week. This is why we added the internet search feature in Part 1, to give it access to current information.

Second, sometimes LLMs can "hallucinate." This means they might confidently make up facts or provide incorrect information, especially when asked about very specific or obscure topics that were not well-represented in their training data. It is not because they are trying to trick us, but because they are trying their best to generate a coherent and plausible response based on patterns they have learned, even if they do not have the exact factual basis.

To make our chatbot truly reliable and an expert on specific subjects, we need a way to provide it with accurate, up-to-date, and domain-specific information before it generates a response. This is where Retrieval Augmented Generation comes to the rescue!


Introducing Retrieval Augmented Generation (RAG)

Think of RAG as a two-step process for your chatbot:

  1. Retrieval: When you ask a question, the chatbot first acts like a super-fast librarian. It quickly scans your personal document library to find all the pieces of information that are most relevant to your question.
  2. Generation: Once it has found those relevant pieces, it then uses its powerful language generation abilities to formulate an answer, but now it has the precise context from your documents right in front of it. It is like having the answer key before taking a test!

This process makes your chatbot more accurate, reduces hallucinations, and allows it to specialize in any knowledge base you provide.

To implement RAG, we will need to learn about a few new concepts and tools, but do not worry, we will go through each one step-by-step.


Setting Up Your RAG Toolkit

Before we start coding, we need to ensure our Python environment has all the necessary libraries. If you have not already, please install these using pip:

pip install langchain langchain-community sentence-transformers faiss-cpu pypdf
  • langchain and langchain-community are our trusty friends for orchestrating LLMs and various components.
  • sentence-transformers helps us turn text into numerical representations.
  • faiss-cpu is a library that helps us store and quickly search these numerical representations.
  • pypdf helps us read PDF documents, which are a common format for knowledge bases.

Step 1: Gathering Your Knowledge Base and Breaking Down Big Ideas (Chunking)

Our chatbot needs documents to learn from. For this tutorial, let us imagine you have a PDF document named my_science_notes.pdf filled with all your amazing science notes. You can replace this with any PDF document you wish to use as your chatbot's knowledge base.

Now, imagine giving a whole textbook to someone and asking them to find one specific answer. It would take a long time! LLMs face a similar problem. They have a limit to how much text they can process at once. So, instead of giving the entire document to the LLM, we break our large documents into smaller, more manageable pieces called "chunks." This process is known as chunking.

Why chunking?

  • Manageable Size: Smaller chunks fit within the LLM's input limit.
  • Relevance: When searching for information, it is easier to find a few relevant small chunks than to sift through a massive document.
  • Efficiency: Processing smaller pieces of text is faster and uses less computational power.

We will use LangChain's RecursiveCharacterTextSplitter to intelligently split our documents. This splitter tries to keep related sentences together by splitting on different characters (like newlines, then spaces, then commas) in a hierarchical way.

Here is how you can load your document and split it into chunks:

from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

# First, specify the path to your document.
# Make sure 'my_science_notes.pdf' is in the same directory as your Python script,
# or provide the full path to the file.
document_path = "my_science_notes.pdf"

# Initialize the PDF loader. This tool knows how to read PDF files.
print(f"Loading document from: {document_path}")
loader = PyPDFLoader(document_path)

# Load the entire document. The loader reads the PDF and extracts its text content.
# It returns a list of 'Document' objects, where each object might represent a page.
documents = loader.load()
print(f"Successfully loaded {len(documents)} pages from the document.")

# Initialize the text splitter. This is like setting up rules for how to break down the text.
# 'chunk_size' determines the maximum number of characters in each piece.
# 'chunk_overlap' specifies how many characters each chunk will share with the previous one.
# Overlapping helps ensure that context is not lost when a key piece of information
# falls exactly on a chunk boundary.
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,  # Each chunk will aim to be around 1000 characters long
    chunk_overlap=200,  # Chunks will overlap by 200 characters to maintain context
    length_function=len, # We are measuring chunk size by character count
)

# Split the loaded documents into smaller, manageable chunks.
print("Splitting documents into smaller chunks...")
chunks = text_splitter.split_documents(documents)
print(f"Original document split into {len(chunks)} chunks.")

# Let us print the first chunk to see what it looks like.
# This helps us understand how the text has been divided.
if chunks:
    print("\n--- Example of a document chunk ---")
    print(chunks[0].page_content)
    print("----------------------------------")
else:
    print("No chunks were created. Please check your document and splitter settings.")

Step 2: Turning Words into Numbers (Embeddings)

Now that we have our document broken into chunks, how does our chatbot find the most relevant chunks for a given question? It cannot just read every chunk and compare it word-for-word. That would be too slow!

This is where embeddings come in. An embedding is a numerical representation of a piece of text. Imagine taking a sentence, a paragraph, or even a whole chunk, and turning it into a long list of numbers (a vector). The magic is that texts with similar meanings will have similar lists of numbers. In a high-dimensional space, these similar numerical lists will be "close" to each other.

So, when you ask a question, we will turn your question into an embedding (a list of numbers). Then, we will compare your question's embedding to the embeddings of all our document chunks. The chunks whose embeddings are numerically "closest" to your question's embedding are the most relevant ones!

We will use a pre-trained embedding model from Hugging Face, specifically sentence-transformers/all-MiniLM-L6-v2, which is excellent for this task and relatively lightweight.

Here is how you generate embeddings for your chunks:

from langchain_community.embeddings import HuggingFaceEmbeddings

# Specify the name of the pre-trained embedding model we want to use.
# This model is designed to convert sentences and paragraphs into meaningful numerical vectors.
embedding_model_name = "sentence-transformers/all-MiniLM-L6-v2"

# Initialize the embedding model.
# The first time you run this, the model will be downloaded to your computer.
# This might take a moment, depending on your internet connection.
print(f"\nInitializing embedding model: {embedding_model_name}")
embeddings = HuggingFaceEmbeddings(model_name=embedding_model_name)
print("Embedding model initialized successfully.")

# You can test the embedding model by converting a simple sentence into a vector.
# This helps confirm that the model is working as expected.
sample_text = "The quick brown fox jumps over the lazy dog."
sample_vector = embeddings.embed_query(sample_text)
print(f"Example: Embedding vector for '{sample_text}' has length: {len(sample_vector)}")
# The length of the vector indicates the dimensionality of the embedding space.

Step 3: Storing and Finding Knowledge (Vector Database)

We have our chunks, and we know how to turn them into numerical embeddings. Now, we need a place to store all these embeddings efficiently and a way to quickly search through them to find the closest ones to our question's embedding. This special storage and search system is called a vector database (or vector store).

For our tutorial, we will use FAISS (Facebook AI Similarity Search). FAISS is not a full-fledged database like SQL or NoSQL, but it is an incredibly fast and efficient library specifically designed for similarity search on large collections of vectors. It is perfect for our needs because it is easy to set up locally and works entirely in memory (or can be saved to disk).

Here is how you create a FAISS vector store from your chunks and embeddings, and how you can save it for future use:

from langchain_community.vectorstores import FAISS

# Ensure 'chunks' and 'embeddings' from the previous steps are available.
# If you are running this code separately, you would need to re-run the chunking
# and embedding model initialization steps first.

# Create a FAISS vector store.
# This step takes all your document chunks, converts each one into an embedding
# using the 'embeddings' model we initialized, and then builds an efficient
# index within FAISS for fast similarity searching.
print("\nCreating FAISS vector store from document chunks...")
vector_store = FAISS.from_documents(chunks, embeddings)
print("FAISS vector store created successfully.")

# It is a good practice to save your vector store to disk.
# This way, you do not have to re-process all your documents and generate embeddings
# every single time you start your chatbot. You can just load the saved index.
vector_store_path = "faiss_index_chatbot"
print(f"Saving vector store to local disk at: {vector_store_path}")
vector_store.save_local(vector_store_path)
print("Vector store saved successfully. You can now load it directly next time.")

# To demonstrate how to load it back, here is the code you would use
# in a new session or after restarting your application:
print(f"\nDemonstrating how to load the vector store from: {vector_store_path}")
# When loading, you must provide the embedding model again, as FAISS needs it
# to understand how to interpret and compare vectors.
loaded_vector_store = FAISS.load_local(
    vector_store_path,
    embeddings,
    # This parameter is important for security; it allows loading objects
    # that might contain custom Python code. For our local use, it is fine,
    # but be cautious with untrusted sources.
    allow_dangerous_deserialization=True
)
print("Vector store loaded successfully for demonstration.")

# You can test the loaded vector store by performing a simple similarity search.
# This will find chunks that are most similar to your query.
test_query = "What are the main components of a cell?"
print(f"\nSearching for documents related to: '{test_query}'")
retrieved_docs = loaded_vector_store.similarity_search(test_query, k=2) # Retrieve top 2 similar documents
print("Retrieved documents (showing content of the first one):")
if retrieved_docs:
    print(retrieved_docs[0].page_content)
else:
    print("No documents retrieved for the test query.")

Step 4: Bringing it All Together: RAG in Action with Your Chatbot

Now for the exciting part: integrating our RAG system into the chatbot we built in Part 1! Our goal is to make sure that when a user asks a question, our chatbot first retrieves relevant information from our FAISS vector store and then uses that information to formulate a more informed answer.

Since our Part 1 chatbot already had conversational memory, we will use LangChain's ConversationalRetrievalChain. This chain is specifically designed to handle both retrieval and chat history, making our chatbot even smarter.

Here is how you would modify your existing chatbot code (from Part 1) to incorporate RAG:

from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain_community.llms import HuggingFacePipeline
from transformers import pipeline
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
import streamlit as st # Assuming your chatbot UI is built with Streamlit

# --- 1. Load the Language Model (LLM) from Part 1 ---
# This part should be similar to how you set up your LLM in Part 1.
# We are using a Hugging Face model (like GPT-2) via a pipeline.
print("\nInitializing the Large Language Model (LLM) pipeline...")
llm_pipeline = pipeline(
    "text-generation",
    model="gpt2",  # Replace with the specific model you used in Part 1 if different
    tokenizer="gpt2",
    max_new_tokens=500,  # Maximum number of tokens the LLM will generate in its response
    device=-1,  # -1 for CPU, 0 for GPU if you have one and PyTorch is configured for it
)
llm = HuggingFacePipeline(pipeline=llm_pipeline)
print("LLM pipeline initialized.")

# --- 2. Load the Embedding Model ---
# We need the same embedding model to load our vector store and to embed new queries.
print("Initializing embedding model for loading vector store...")
embedding_model_name = "sentence-transformers/all-MiniLM-L6-v2"
embeddings = HuggingFaceEmbeddings(model_name=embedding_model_name)
print("Embedding model initialized.")

# --- 3. Load Your FAISS Vector Store ---
# This loads the knowledge base you created and saved in the previous step.
vector_store_path = "faiss_index_chatbot"
print(f"Loading FAISS vector store from: {vector_store_path}")
try:
    loaded_vector_store = FAISS.load_local(
        vector_store_path,
        embeddings,
        allow_dangerous_deserialization=True
    )
    print("Vector store loaded successfully.")
except Exception as e:
    print(f"Error loading vector store: {e}")
    print("Please ensure you have run Step 3 to create and save the 'faiss_index_chatbot'.")
    # Handle the error, perhaps by exiting or using a fallback mechanism
    st.error("Could not load knowledge base. Please ensure it is created.")
    st.stop() # Stop Streamlit execution if the knowledge base isn't available

# --- 4. Create a Retriever from the Vector Store ---
# The retriever is the component that knows how to search your vector store
# for relevant documents based on a query.
# 'k=4' means it will retrieve the top 4 most relevant document chunks.
retriever = loaded_vector_store.as_retriever(search_kwargs={"k": 4})
print("Retriever created from the vector store.")

# --- 5. Initialize Conversational Memory (from Part 1) ---
# This memory object will keep track of the conversation history,
# allowing the chatbot to remember previous turns.
print("Initializing conversational memory...")
memory = ConversationBufferMemory(
    memory_key="chat_history",  # This key tells LangChain where to find the chat history
    return_messages=True        # We want the memory to return actual message objects
)
print("Conversational memory initialized.")

# --- 6. Create the Conversational RAG Chain ---
# This is the core of our RAG chatbot. It combines the LLM, the retriever,
# and the conversational memory into a single, powerful chain.
# 'condense_question_llm=llm' means the LLM itself will be used to
# rephrase the user's question, taking into account the chat history,
# to make it a standalone question for better retrieval.
# 'return_source_documents=True' is very useful for debugging and for
# showing the user where the information came from.
print("Creating the Conversational Retrieval Chain...")
qa_chain = ConversationalRetrievalChain.from_llm(
    llm=llm,
    retriever=retriever,
    memory=memory,
    condense_question_llm=llm, # Use the same LLM to condense questions
    return_source_documents=True
)
print("Conversational Retrieval Chain created. Your RAG chatbot is ready!")

# --- 7. Integrate into your Streamlit App (Conceptual Example) ---
# This section shows how you would typically integrate 'qa_chain' into your
# existing Streamlit application's chat loop.

# Assuming your Streamlit app has a way to get user input and display responses:

# if "messages" not in st.session_state:
#     st.session_state.messages = []

# for message in st.session_state.messages:
#     with st.chat_message(message["role"]):
#         st.markdown(message["content"])

# user_query = st.chat_input("Ask your RAG-powered chatbot:")

# if user_query:
#     st.session_state.messages.append({"role": "user", "content": user_query})
#     with st.chat_message("user"):
#         st.markdown(user_query)

#     with st.chat_message("assistant"):
#         # Instead of calling your old conversation chain, call the new qa_chain
#         # The input key for ConversationalRetrievalChain is "question"
#         response = qa_chain.invoke({"question": user_query})
#         bot_answer = response["answer"]
#         st.markdown(bot_answer)

#         # Optionally, display the source documents
#         if response["source_documents"]:
#             st.subheader("Sources:")
#             for i, doc in enumerate(response["source_documents"]):
#                 st.text(f"Source {i+1}: {doc.metadata.get('source', 'Unknown source')}")
#                 st.text(doc.page_content[:200] + "...") # Show first 200 chars of source
#         st.session_state.messages.append({"role": "assistant", "content": bot_answer})

print("\nConceptual Streamlit integration shown. You will replace your old chatbot call with 'qa_chain.invoke'.")
print("Remember to adapt your Streamlit UI to display source documents if desired.")

In the code above, we first load our LLM and the embedding model, just like before. Then, we load our pre-built FAISS vector store. We turn this vector store into a retriever, which is LangChain's way of saying "the thing that fetches documents." We also bring back our ConversationBufferMemory to keep our chatbot conversational.

Finally, we create the ConversationalRetrievalChain. When you ask a question, this chain will:

  1. Look at your current question and the chat history to form a clear, standalone question.
  2. Use this standalone question to query the retriever and find the most relevant chunks from your faiss_index_chatbot.
  3. Combine these retrieved chunks with the chat history and your original question, and send all of this context to the LLM.
  4. The LLM then generates an answer based on all this rich information.

The return_source_documents=True parameter is incredibly useful because it makes the chain return not just the answer but also the specific document chunks it used to generate that answer. This helps you (and your users) verify the information and understand its origin.


Step 5: Testing and Refining Your RAG Chatbot

Now that your chatbot is equipped with RAG, it is time to test it out!

Run your Streamlit application and try asking questions that are specifically related to the content of your my_science_notes.pdf document.

For example, if your science notes are about photosynthesis, ask: "What are the main stages of photosynthesis?" or "What role does chlorophyll play?"

Compare the answers you get now with the answers you might have received from the chatbot in Part 1 (without RAG). You should notice that the RAG-powered chatbot provides more precise, detailed, and accurate answers directly from your document.

You might also want to experiment with:

  • chunk_size and chunk_overlap: How do different chunking strategies affect the quality of retrieved documents and answers?
  • k in retriever.as_retriever(search_kwargs={"k": 4}): What happens if you retrieve more or fewer documents? Does retrieving too many confuse the LLM, or does retrieving too few miss important context?
  • The content of your documents: The quality of your RAG system heavily depends on the quality and comprehensiveness of your knowledge base.

Remember, building AI applications is an iterative process. You will continuously learn, experiment, and refine your creation.


Conclusion: Your Smart, Knowledgeable Chatbot!

Congratulations! You have successfully extended your basic AI chatbot into a powerful, knowledge-augmented system using Retrieval Augmented Generation. You have learned how to:

  • Break down large documents into manageable chunks.
  • Convert text into numerical embeddings that capture meaning.
  • Store and efficiently search these embeddings using a vector database like FAISS.
  • Integrate this retrieval capability with your LLM and conversational memory using LangChain.

Your chatbot is now not just a general conversationalist but also a specialized expert, capable of drawing information directly from your custom knowledge base. This opens up a world of possibilities for creating AI assistants tailored to specific subjects, industries, or even just your personal learning needs.

Keep experimenting, keep learning, and keep building! The world of AI is vast and exciting, and you are now equipped with even more powerful tools to explore it. What will you teach your chatbot next? The possibilities are truly endless!



The Complete RAG Chatbot Application Code

Here is the full Python script for your RAG-powered Streamlit chatbot. This code combines all the steps we covered: loading documents, chunking, creating embeddings, building a FAISS vector store, and integrating it with a conversational LLM and memory.

import streamlit as st
import os
import torch
from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM

from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain_community.llms import HuggingFacePipeline
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

# --- Configuration ---
# Define the path for your knowledge base PDF file.
# IMPORTANT: Place your 'my_science_notes.pdf' file in the same directory as this script.
# If the file doesn't exist, the script will prompt you to create it.
KNOWLEDGE_BASE_PDF_PATH = "my_science_notes.pdf"

# Define the path where your FAISS vector store will be saved/loaded.
FAISS_INDEX_PATH = "faiss_index_chatbot"

# Define the Hugging Face model for embeddings.
EMBEDDING_MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"

# Define the Hugging Face model for the Large Language Model (LLM).
# This should be the same model you used in Part 1.
# Examples: "gpt2", "distilgpt2", "microsoft/DialoGPT-small"
LLM_MODEL_NAME = "gpt2"

# --- Helper Function to Initialize Knowledge Base ---
def initialize_knowledge_base():
    """
    Loads documents, chunks them, creates embeddings, and builds/saves a FAISS vector store.
    If the FAISS index already exists, it loads it.
    """
    # Initialize embeddings model first, as it's needed for both creating and loading the FAISS index.
    print(f"Initializing embedding model: {EMBEDDING_MODEL_NAME}")
    embeddings = HuggingFaceEmbeddings(model_name=EMBEDDING_MODEL_NAME)
    print("Embedding model initialized.")

    if os.path.exists(FAISS_INDEX_PATH):
        print(f"Loading existing FAISS vector store from: {FAISS_INDEX_PATH}")
        try:
            # When loading, we must provide the embedding model again, as FAISS needs it
            # to understand how to interpret and compare vectors.
            vector_store = FAISS.load_local(FAISS_INDEX_PATH, embeddings, allow_dangerous_deserialization=True)
            print("Vector store loaded successfully.")
            return vector_store
        except Exception as e:
            st.error(f"Error loading existing vector store: {e}. Attempting to rebuild.")
            # If loading fails, proceed to rebuild
            pass
    
    print("FAISS vector store not found or failed to load. Initializing from documents...")
    if not os.path.exists(KNOWLEDGE_BASE_PDF_PATH):
        st.error(f"Error: Knowledge base PDF '{KNOWLEDGE_BASE_PDF_PATH}' not found.")
        st.info("Please create a PDF file named 'my_science_notes.pdf' in the same directory as this script, "
                "filled with information you want your chatbot to learn. For example, you could put notes about "
                "biology, chemistry, or physics in it.")
        st.stop() # Stop the Streamlit app if the PDF is missing

    print(f"Loading document from: {KNOWLEDGE_BASE_PDF_PATH}")
    # Initialize the PDF loader. This tool knows how to read PDF files.
    loader = PyPDFLoader(KNOWLEDGE_BASE_PDF_PATH)
    # Load the entire document. The loader reads the PDF and extracts its text content.
    # It returns a list of 'Document' objects, where each object might represent a page.
    documents = loader.load()
    print(f"Successfully loaded {len(documents)} pages from the document.")

    # Initialize the text splitter. This is like setting up rules for how to break down the text.
    # 'chunk_size' determines the maximum number of characters in each piece.
    # 'chunk_overlap' specifies how many characters each chunk will share with the previous one.
    # Overlapping helps ensure that context is not lost when a key piece of information
    # falls exactly on a chunk boundary.
    text_splitter = RecursiveCharacterTextSplitter(
        chunk_size=1000,  # Each chunk will aim to be around 1000 characters long
        chunk_overlap=200,  # Chunks will overlap by 200 characters to maintain context
        length_function=len, # We are measuring chunk size by character count
    )
    # Split the loaded documents into smaller, manageable chunks.
    chunks = text_splitter.split_documents(documents)
    print(f"Original document split into {len(chunks)} chunks.")

    print("Creating FAISS vector store from document chunks...")
    # This step takes all your document chunks, converts each one into an embedding
    # using the 'embeddings' model we initialized, and then builds an efficient
    # index within FAISS for fast similarity searching.
    vector_store = FAISS.from_documents(chunks, embeddings)
    print("FAISS vector store created successfully.")

    print(f"Saving vector store to local disk at: {FAISS_INDEX_PATH}")
    # It is a good practice to save your vector store to disk.
    # This way, you do not have to re-process all your documents and generate embeddings
    # every single time you start your chatbot. You can just load the saved index.
    vector_store.save_local(FAISS_INDEX_PATH)
    print("Vector store saved successfully.")
    return vector_store

# --- Streamlit UI Setup ---
# Configure the Streamlit page with a title, icon, and layout.
st.set_page_config(page_title="RAG-Powered Chatbot", page_icon="📚", layout="wide")
st.title("📚 Your RAG-Powered Chatbot (Part 2)")
st.write("Hello, future AI developers! This chatbot now uses Retrieval Augmented Generation (RAG) to answer questions based on your custom knowledge base!")
st.markdown("---")

# --- Initialize Components (using st.session_state to avoid re-running on every interaction) ---
# Streamlit's `st.session_state` is crucial here. It allows us to store objects
# (like our LLM, vector store, and chains) so they are not re-initialized
# every time the user interacts with the app (e.g., typing a new message).

# Initialize LLM (Large Language Model)
if "llm" not in st.session_state:
    print(f"\nInitializing LLM pipeline: {LLM_MODEL_NAME}...")
    try:
        # Load the tokenizer and model for the specified LLM.
        tokenizer = AutoTokenizer.from_pretrained(LLM_MODEL_NAME)
        model = AutoModelForCausalLM.from_pretrained(LLM_MODEL_NAME)
        
        # Determine the device to run the LLM on (GPU if available, otherwise CPU).
        device = 0 if torch.cuda.is_available() else -1
        print(f"Using device: {'cuda' if device == 0 else 'cpu'}")

        # Create a Hugging Face pipeline for text generation.
        pipe = pipeline(
            "text-generation",
            model=model,
            tokenizer=tokenizer,
            max_new_tokens=500,  # Maximum number of tokens the LLM will generate in its response
            device=device,       # Assign the model to the detected device
            do_sample=True,      # Enable sampling for more varied responses
            temperature=0.7,     # Control creativity (lower for more deterministic, higher for more creative)
            top_k=50,            # Limit sampling to top_k most likely words
            top_p=0.95           # Limit sampling to words that make up top_p probability mass
        )
        # Wrap the Hugging Face pipeline in LangChain's HuggingFacePipeline class.
        st.session_state.llm = HuggingFacePipeline(pipeline=pipe)
        print("LLM pipeline initialized.")
    except Exception as e:
        st.error(f"Failed to load LLM ({LLM_MODEL_NAME}). Please ensure it's installed and accessible. "
                 f"You might need to run `pip install transformers torch` and check your internet connection. Error: {e}")
        st.stop() # Stop the app if the LLM cannot be loaded.

# Initialize Vector Store (Knowledge Base)
if "vector_store" not in st.session_state:
    st.session_state.vector_store = initialize_knowledge_base()

# Create Retriever
if "retriever" not in st.session_state:
    # The retriever is the component that knows how to search your vector store
    # for relevant documents based on a query.
    # 'k=4' means it will retrieve the top 4 most relevant document chunks.
    st.session_state.retriever = st.session_state.vector_store.as_retriever(search_kwargs={"k": 4})
    print("Retriever created from the vector store.")

# Initialize Conversational Memory
if "memory" not in st.session_state:
    # This memory object will keep track of the conversation history,
    # allowing the chatbot to remember previous turns.
    st.session_state.memory = ConversationBufferMemory(
        memory_key="chat_history",  # This key tells LangChain where to find the chat history
        return_messages=True        # We want the memory to return actual message objects
    )
    print("Conversational memory initialized.")

# Create Conversational Retrieval Chain
if "qa_chain" not in st.session_state:
    # This is the core of our RAG chatbot. It combines the LLM, the retriever,
    # and the conversational memory into a single, powerful chain.
    # 'condense_question_llm=st.session_state.llm' means the LLM itself will be used to
    # rephrase the user's question, taking into account the chat history,
    # to make it a standalone question for better retrieval.
    # 'return_source_documents=True' is very useful for debugging and for
    # showing the user where the information came from.
    print("Creating the Conversational Retrieval Chain...")
    st.session_state.qa_chain = ConversationalRetrievalChain.from_llm(
        llm=st.session_state.llm,
        retriever=st.session_state.retriever,
        memory=st.session_state.memory,
        condense_question_llm=st.session_state.llm, # Use the same LLM to condense questions
        return_source_documents=True
    )
    print("Conversational Retrieval Chain created. Your RAG chatbot is ready!")

# --- Display Chat History ---
# If there are no messages in the session state, initialize an empty list.
if "messages" not in st.session_state:
    st.session_state.messages = []

# Iterate through all stored messages and display them in the chat interface.
for message in st.session_state.messages:
    with st.chat_message(message["role"]):
        st.markdown(message["content"])
        # If a message has associated sources, display them in an expandable section.
        if "sources" in message and message["sources"]:
            with st.expander("Sources Used"):
                for i, source in enumerate(message["sources"]):
                    st.text(f"Source {i+1}: {source}")

# --- Handle User Input ---
# Create a text input box for the user to type their questions.
user_query = st.chat_input("Ask your RAG-powered chatbot a question about your knowledge base!")

# If the user has entered a query:
if user_query:
    # Add the user's query to the chat history.
    st.session_state.messages.append({"role": "user", "content": user_query})
    with st.chat_message("user"):
        st.markdown(user_query)

    # Display a "Thinking..." spinner while the chatbot processes the query.
    with st.chat_message("assistant"):
        with st.spinner("Thinking..."):
            try:
                # Call the RAG-powered conversational chain to get a response.
                response = st.session_state.qa_chain.invoke({"question": user_query})
                bot_answer = response["answer"]
                source_documents = response.get("source_documents", [])

                # Display the chatbot's answer.
                st.markdown(bot_answer)

                # Display source documents if available.
                if source_documents:
                    with st.expander("Sources Used"):
                        for i, doc in enumerate(source_documents):
                            # Extract filename from the full path for cleaner display.
                            file_name = os.path.basename(doc.metadata.get('source', 'Unknown file'))
                            st.markdown(f"**Source {i+1}** (Page: {doc.metadata.get('page', 'N/A')}, "
                                        f"File: {file_name}):")
                            st.text(doc.page_content[:300] + "...") # Show first 300 characters of source content
                
                # Store the assistant's message and sources in session state for display in future turns.
                st.session_state.messages.append({"role": "assistant", "content": bot_answer,
                                                   "sources": [f"Page {doc.metadata.get('page', 'N/A')} from "
                                                               f"{os.path.basename(doc.metadata.get('source', 'Unknown file'))}"
                                                               for doc in source_documents]})
            except Exception as e:
                # Handle any errors that occur during the RAG process.
                st.error(f"An error occurred while processing your request: {e}")
                st.session_state.messages.append({"role": "assistant", "content": f"I'm sorry, I encountered an error: {e}"})

How to Build, Deploy, and Start Your RAG Chatbot Application

Now that you have the complete code, let's get it up and running! This process involves a few straightforward steps:

Step 1: Prepare Your Environment (Install Dependencies)

First, you need to make sure your Python environment has all the necessary libraries installed. Open your terminal or command prompt and run the following command:

pip install langchain langchain-community sentence-transformers faiss-cpu pypdf transformers torch streamlit
  • langchain and langchain-community: These are the core libraries for building LLM applications, helping us orchestrate all the components.
  • sentence-transformers: This library provides the embedding model we use to convert text into numerical vectors.
  • faiss-cpu: This is Facebook AI Similarity Search, a super-fast library for storing and searching our text embeddings.
  • pypdf: This library allows our application to read and extract text from PDF documents.
  • transformers and torch: These are essential for loading and running the Hugging Face Large Language Model (LLM) you chose (like GPT-2). torch is the underlying deep learning framework.
  • streamlit: This is the framework we use to create the interactive web-based user interface for our chatbot.

Step 2: Save the Application Code

  1. Create a new file: Open a plain text editor (like VS Code, Sublime Text, Notepad++, or even a simple text editor).
  2. Copy the code: Copy the entire Python script provided above.
  3. Save the file: Save the copied code into a file named rag_chatbot_app.py (or any other .py name you prefer) in a folder on your computer. Remember the location of this folder!

Step 3: Create Your Knowledge Base Document

Your chatbot needs something to learn from!

  1. Create a PDF file: In the same folder where you saved rag_chatbot_app.py, create a PDF document named my_science_notes.pdf.
    • Content: Fill this PDF with information you want your chatbot to be an expert on. For example, you could:

      • Copy-paste notes from your science textbook about topics like photosynthesis, the solar system, or human anatomy.
      • Write a short story or a fictional history.
      • Include details about a specific project or hobby.
    • How to create a PDF: You can use a word processor (like Microsoft Word, Google Docs, LibreOffice Writer) to type out your content and then use its "Save As PDF" or "Print to PDF" function. Make sure the saved file is named my_science_notes.pdf.

    • Example Content for my_science_notes.pdf (you can copy this into a document and save as PDF):

      The process of photosynthesis is how plants convert light energy into chemical energy,
      which is stored in glucose. This process primarily occurs in the chloroplasts,
      specifically using chlorophyll, the green pigment. Photosynthesis requires
      carbon dioxide, water, and sunlight. It produces glucose (sugar) and oxygen.
      The main stages are the light-dependent reactions and the light-independent reactions (Calvin Cycle).
      During light-dependent reactions, light energy is captured by chlorophyll and converted
      into ATP and NADPH, releasing oxygen as a byproduct. The Calvin Cycle then uses
      this ATP and NADPH to convert carbon dioxide into glucose.
      
      The human heart is a muscular organ that pumps blood throughout the body.
      It has four chambers: two atria and two ventricles. The right side of the heart
      pumps deoxygenated blood to the lungs, while the left side pumps oxygenated blood
      to the rest of the body. The average adult heart beats about 60 to 100 times per minute.
      

Step 4: Run Your Streamlit Application

Now for the moment of truth!

  1. Open your terminal/command prompt: Navigate to the folder where you saved rag_chatbot_app.py and my_science_notes.pdf. You can do this using the cd command (e.g., cd path/to/your/folder).

  2. Start the Streamlit app: Once you are in the correct directory, run the following command:

    streamlit run rag_chatbot_app.py
    
    • First Run: The first time you run this, it might take a few minutes. Streamlit will start, and the application will:
      • Download the Hugging Face gpt2 model (if not already cached).
      • Download the sentence-transformers/all-MiniLM-L6-v2 embedding model.
      • Load your my_science_notes.pdf, chunk it, create embeddings, and build the FAISS vector store. This process will print messages to your terminal indicating its progress.
      • Finally, it will save the FAISS index to a folder named faiss_index_chatbot in your application directory.
  3. Access the application: Once it finishes loading, your web browser should automatically open a new tab displaying your RAG-powered chatbot. If it does not, look for a URL in your terminal output that starts with http://localhost:8501 and open it manually.

Step 5: Interact with Your Chatbot!

You now have a fully functional RAG chatbot!

  1. Ask questions from your PDF: Type questions into the chat input box that are specifically related to the content you put in my_science_notes.pdf.
    • Try questions like: "What is photosynthesis and where does it occur?" or "What are the main stages of photosynthesis?"
    • Or, if you included information about the heart: "How many chambers does the human heart have?" or "What is the function of the right side of the heart?"
  2. Observe the responses: You should see your chatbot provide answers that are directly informed by the content of your PDF.
  3. Check the sources: Below the chatbot's answer, you will find an expandable section labeled "Sources Used." Click on it to see which specific chunks from your PDF the chatbot used to formulate its response. This is a powerful feature for verifying information!
  4. Test general knowledge: You can also ask general questions not found in your PDF. The LLM will still try to answer these based on its pre-trained knowledge.

What Happens Next Time You Run the App?

Because we added the logic to save and load the FAISS index, when you run streamlit run rag_chatbot_app.pyagain, it will detect the faiss_index_chatbot folder, load the pre-built vector store, and skip the time-consuming process of re-chunking and re-embedding your PDF. This makes subsequent startups much faster!

Congratulations! You have successfully built and deployed a sophisticated RAG-powered chatbot. This is a significant step in your AI journey, giving you the power to create intelligent agents that are experts in specific domains. Keep experimenting, keep learning, and enjoy your smarter chatbot!

Saturday, October 10, 2026

BUILDING YOUR OWN LLM CHATBOT IN PYTHON - PART 1/2: FROM FIRST STEPS TO WORKING IMPLEMENTATION


Part 1/2

INTRODUCTION: WELCOME TO THE WORLD OF AI CHATBOTS

Imagine having a conversation with a computer program that understands you, remembers what you talked about earlier, and can even search the internet to give you up-to-date information. This is exactly what we are going to build together in this tutorial. We will create a fully functional chatbot powered by Large Language Models, or LLMs for short.

Large Language Models are artificial intelligence systems that have been trained on massive amounts of text from books, websites, and other sources. They can understand human language and generate responses that sound remarkably human-like. You have probably used chatbots like ChatGPT or Claude before. In this tutorial, you will learn how to build your own version from scratch.

What makes this tutorial special is that we will build a chatbot that has real, practical features. Our chatbot will remember previous conversations, so you can have natural back-and-forth discussions just like texting with a friend. It will save different chat sessions so you can come back to old conversations later. Most excitingly, it will have a special internet search mode that allows it to look up current information online when you activate it with an icon, giving you the best of both worlds - the knowledge in its training and fresh information from the web.

We will also make sure our chatbot works efficiently on different types of computers. Whether you have a Mac with Apple Silicon chips or a Windows or Linux computer with an NVIDIA graphics card, our code will automatically detect and use the best available hardware to run the AI model quickly.

This tutorial is designed for high school students who have some basic Python programming knowledge. You should be comfortable with variables, functions, and simple data structures like lists and dictionaries. Do not worry if you have never worked with AI before - we will explain everything step by step, and every piece of code will be fully explained so you understand exactly what it does and why.


UNDERSTANDING THE TECHNOLOGY STACK

Before we start coding, let us understand the main tools and frameworks we will use. Think of building software like constructing a house - you need different materials and tools for different parts. Our chatbot will use several specialized tools, each designed for a specific purpose.

The first major tool is Streamlit. Streamlit is a Python framework that makes it incredibly easy to create web applications. Normally, building a web interface requires knowledge of HTML, CSS, and JavaScript. With Streamlit, you can create beautiful web interfaces using only Python. This is perfect for us because we can focus on the chatbot logic without getting bogged down in web development details. Streamlit will handle creating the chat window, displaying messages, and managing user input.

Next, we have LangChain, which is a framework specifically designed for building applications with Large Language Models. LangChain provides pre-built components for common tasks like managing conversation history, calling different AI models, and integrating external tools like web search. Instead of writing hundreds of lines of code from scratch, we can use LangChain’s ready-made components and focus on making our chatbot unique and useful.

For the actual AI brain of our chatbot, we will use the Transformers library from Hugging Face. Hugging Face is a company that has made thousands of pre-trained AI models available to developers for free. The Transformers library makes it easy to download these models and use them in your own applications. We will use a relatively small model that can run on regular computers rather than requiring expensive cloud servers.

For internet search capabilities, we will use DuckDuckGo Search. DuckDuckGo is a privacy-focused search engine, and LangChain provides a tool that lets our chatbot perform web searches. When you activate search mode, the chatbot will query DuckDuckGo and use the search results to give you current information.

Finally, PyTorch is the deep learning framework that powers the AI model. PyTorch includes support for different types of hardware acceleration. On Apple computers with M{1,2,3,4,5,6} chips, it can use Metal Performance Shaders or MPS for short to run the model faster. On computers with NVIDIA graphics cards, it uses CUDA. Our code will automatically detect what hardware you have and use it optimally.


SETTING UP YOUR DEVELOPMENT ENVIRONMENT

The first step in any programming project is setting up your development environment. This means installing Python and all the necessary libraries. We will walk through this process step by step.

First, make sure you have Python installed on your computer. You need Python version three point eight or higher. You can check if Python is installed by opening your terminal or command prompt and typing “python –version”. If Python is not installed, download it from python.org and follow the installation instructions for your operating system.

Once Python is installed, we strongly recommend creating a virtual environment. A virtual environment is like a separate workspace for your project that keeps all your project’s libraries isolated from other Python projects on your computer. This prevents conflicts between different versions of libraries. To create a virtual environment, open your terminal, navigate to the folder where you want to keep your project, and run these commands. On Windows, type “python -m venv chatbot_env” to create the environment, then “chatbot_env\Scripts\activate” to activate it. On Mac or Linux, type “python3 -m venv chatbot_env” to create it, then “source chatbot_env/bin/activate” to activate it.

When your virtual environment is activated, you will see its name in parentheses at the start of your command prompt. This tells you that any packages you install will go into this isolated environment rather than affecting your system-wide Python installation.


INSTALLING REQUIRED PACKAGES

Now we need to install all the Python packages our chatbot will use. We will install these one by one so you understand what each package does.

First, install Streamlit by typing “pip install streamlit”. This installs the web framework we will use to build our user interface.

Next, install LangChain and its community packages with “pip install langchain langchain-community”. LangChain Community contains additional integrations including the DuckDuckGo search tool.

Now install the Transformers library from Hugging Face with “pip install transformers”. This gives us access to pre-trained language models.

Install PyTorch, which is the deep learning framework that runs the models. The installation command depends on your system. For Mac with Apple Silicon, type “pip install torch torchvision torchaudio”. For Windows or Linux with NVIDIA GPU, visit pytorch.org to get the specific command for your CUDA version. For CPU-only systems, type “pip install torch torchvision torchaudio –index-url https://download.pytorch.org/whl/cpu”.

Install the Accelerate library with “pip install accelerate”. This library helps optimize model loading and inference across different hardware.

Install the DuckDuckGo search tool with “pip install duckduckgo-search”. This enables web search functionality.

Finally, install the Sentence Transformers library with “pip install sentence-transformers”. This provides efficient text embedding models that we might use for future enhancements.

If you want to install all these packages at once, you can create a file called “requirements.txt” with all the package names, then run “pip install -r requirements.txt”. Here is what your requirements.txt file should contain:


streamlit==1.29.0

langchain==0.1.0

langchain-community==0.0.13

transformers==4.36.0

torch==2.1.0

accelerate==0.25.0

duckduckgo-search==4.1.0

sentence-transformers==2.2.2


UNDERSTANDING HOW OUR CHATBOT WILL WORK

Before we start writing code, let us understand the architecture of our chatbot. Architecture means the overall structure and design of how different parts work together.

Our chatbot will have three main layers. The first layer is the user interface layer, which is what you see and interact with. This includes the chat window where messages appear, the text input box where you type your questions, and the search icon that toggles internet search mode. Streamlit handles all of this for us.

The second layer is the logic layer, which contains the brain of our application. This layer manages conversation history, decides when to use internet search, and coordinates between different components. This is where most of our custom code will live.

The third layer is the model layer, which contains the actual AI model that generates responses. This layer also includes the search tool for when internet lookup is needed. LangChain and Transformers libraries handle this layer.

Our chatbot will work like this. When you type a message and press enter, the user interface layer captures your input. The logic layer checks if search mode is activated. If search mode is off, the chatbot uses only the AI model to generate a response based on its training data and the conversation history. If search mode is on, the logic layer first performs a web search using DuckDuckGo, then provides both your question and the search results to the AI model, which formulates a comprehensive answer.

The conversation history is stored in memory during your session. Each session gets a unique identifier, and we save the history to a file when you end the session. This way, you can resume conversations later.


STEP ONE: CREATING THE BASIC STREAMLIT INTERFACE

Let us start by creating a simple Streamlit interface. This will be the foundation that we build upon. Create a new file called “chatbot_app.py” and open it in your text editor.

At the top of the file, we import the necessary libraries:


import streamlit as st

import torch

from datetime import datetime

import json

import os


Now let us configure the Streamlit page. This sets the title that appears in your browser tab and the layout:


st.set_page_config(

    page_title="My AI Chatbot",

    page_icon="🤖",

    layout="wide"

)


Next, we add a title to our page:


st.title("🤖 My AI Chatbot")

st.markdown("Chat with an AI powered by Large Language Models")


Streamlit uses something called session state to remember information between interactions. When a user clicks a button or types something, Streamlit reruns your entire script. Session state lets us preserve data across these reruns. We will initialize our session state variables:


# Initialize session state variables

if 'messages' not in st.session_state:

    st.session_state.messages = []


if 'search_enabled' not in st.session_state:

    st.session_state.search_enabled = False


if 'session_id' not in st.session_state:

    st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")


The messages list will store all the conversation messages. The search_enabled flag tells us whether internet search is currently activated. The session_id is a unique identifier for this conversation session based on the current date and time.

Now let us create the chat interface. Streamlit provides a special container for displaying chat messages:


# Display chat messages

for message in st.session_state.messages:

    with st.chat_message(message["role"]):

        st.markdown(message["content"])


This loop goes through all our saved messages and displays them. Each message has a role, which is either “user” for messages you typed or “assistant” for responses from the chatbot. The chat_message function displays the message with an appropriate icon.

Finally, we add an input box where users can type their messages:


# Chat input

if prompt := st.chat_input("Type your message here..."):

    # Add user message to chat

    st.session_state.messages.append({"role": "user", "content": prompt})

    with st.chat_message("user"):

        st.markdown(prompt)

    

    # For now, just echo the message back

    response = f"You said: {prompt}"

    st.session_state.messages.append({"role": "assistant", "content": response})

    with st.chat_message("assistant"):

        st.markdown(response)


The walrus operator “:=” assigns the user’s input to the prompt variable and checks if it is not empty, all in one line. If the user typed something, we add it to our messages list, display it, create a simple response, and display that too.

You can test this basic interface by saving the file and running “streamlit run chatbot_app.py” in your terminal. A web browser should open showing your chatbot interface. It will not be very smart yet - it just echoes back what you type - but you can see the chat interface working.


STEP TWO: ADDING HARDWARE DETECTION FOR GPU SUPPORT

Before we integrate the AI model, we need to write code that automatically detects what hardware acceleration is available. This is important because it ensures our chatbot runs as fast as possible on any computer.

Add this function near the top of your file, after the imports:


def get_device():

    """

    Detect the best available device for running the model.

    Returns a torch device object.

    """

    # Check for CUDA (NVIDIA GPU)

    if torch.cuda.is_available():

        device = torch.device("cuda")

        device_name = torch.cuda.get_device_name(0)

        print(f"Using CUDA device: {device_name}")

        return device

    

    # Check for MPS (Apple Silicon)

    elif torch.backends.mps.is_available():

        device = torch.device("mps")

        print("Using Apple MPS device")

        return device

    

    # Fall back to CPU

    else:

        device = torch.device("cpu")

        print("Using CPU device")

        return device


This function tries to detect GPU acceleration in order of preference. NVIDIA GPUs with CUDA are generally the fastest for AI models, so we check for that first. If we are on a Mac with Apple Silicon, we check for MPS availability. If neither is available, we fall back to using the CPU, which works on any computer but is slower.

Let us call this function and store the device:


# Get the best available device

device = get_device()


Now our chatbot knows what hardware to use for running the AI model.


STEP THREE: LOADING THE LANGUAGE MODEL

Now comes the exciting part - loading an actual AI language model. We will use a model called GPT-2, which is a smaller model that can run on regular computers. While it is not as powerful as models like GPT-4, it is perfect for learning and can run locally without needing expensive cloud services.

First, let us import the necessary classes from the Transformers library:


from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline


Now we will create a function that loads the model with caching so it only loads once:


@st.cache_resource

def load_model():

    """

    Load the language model and tokenizer.

    Uses Streamlit caching to avoid reloading on every interaction.

    """

    model_name = "gpt2"  # Using GPT-2 as our base model

    

    # Load the tokenizer

    # The tokenizer converts text into numbers the model can understand

    tokenizer = AutoTokenizer.from_pretrained(model_name)

    

    # Load the model

    model = AutoModelForCausalLM.from_pretrained(

        model_name,

        torch_dtype=torch.float16 if device.type != "cpu" else torch.float32

    )

    

    # Move model to the detected device (GPU or CPU)

    model = model.to(device)

    

    # Create a text generation pipeline

    generator = pipeline(

        "text-generation",

        model=model,

        tokenizer=tokenizer,

        device=device.type if device.type != "mps" else -1

    )

    

    return generator, tokenizer


The @st.cache_resource decorator is important. It tells Streamlit to run this function only once and remember the result. Without this, the model would reload every time you send a message, which would be extremely slow.

Let us break down what this function does. First, it specifies the model name. GPT-2 is available in several sizes, and we are using the smallest version. The tokenizer is loaded first - it converts human text into numbers that the model understands. Then we load the actual model.

The torch_dtype parameter specifies the precision of numbers used in the model. Float16 uses half-precision math, which is faster on GPUs and uses less memory. We only use this on GPU though - CPUs work better with float32.

We move the model to our detected device using the to method. Then we create a pipeline, which is a high-level interface that handles all the details of generating text.

Now let us actually load the model:


# Load the model when the app starts

with st.spinner("Loading AI model... This may take a minute..."):

    generator, tokenizer = load_model()


st.success("Model loaded successfully!")


The spinner shows a loading message while the model loads. This is important because loading can take 30 seconds to a minute the first time. The success message confirms everything is ready.


STEP FOUR: CREATING THE CONVERSATION MANAGER

To have natural conversations, we need to manage the conversation history properly. The AI model needs context from previous messages to generate relevant responses. Let us create a class that handles this:


class ConversationManager:

    """

    Manages conversation history and generates responses.

    """

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages):

        """

        Format the conversation history into a prompt for the model.

        """

        # Take only the last max_history messages to avoid exceeding token limits

        recent_messages = messages[-self.max_history:]

        

        # Build the conversation string

        conversation = ""

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        # Add the prompt for the next assistant response

        conversation += "Assistant:"

        

        return conversation

    

    def generate_response(self, messages):

        """

        Generate a response based on conversation history.

        """

        # Format the conversation

        prompt = self.format_conversation(messages)

        

        # Generate response

        outputs = self.generator(

            prompt,

            max_new_tokens=100,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        # Extract the generated text

        generated_text = outputs[0]['generated_text']

        

        # Remove the prompt part to get only the new response

        response = generated_text[len(prompt):].strip()

        

        # Clean up the response

        # Sometimes the model generates multiple lines, we want only the first response

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


This class encapsulates all the logic for managing conversations and generating responses. The format_conversation method creates a prompt string that includes the recent conversation history. We limit this to the last few messages to avoid exceeding the model’s maximum input length.

The generate_response method calls the model to generate a response. Let us understand the parameters. The max_new_tokens parameter limits how long the response can be. The temperature parameter controls randomness - lower values make responses more focused and deterministic, higher values make them more creative. We set do_sample to True to enable random sampling, which makes responses more natural. The pad_token_id parameter handles technical details about how sequences are padded.

After generation, we extract just the new text that the model added and clean it up.

Let us create an instance of our conversation manager:


# Initialize conversation manager

conversation_manager = ConversationManager(generator, tokenizer)


STEP FIVE: ADDING INTERNET SEARCH CAPABILITY

Now let us add the ability to search the internet. This is what makes our chatbot really useful - it can provide up-to-date information instead of being limited to its training data.

First, import the search tool from LangChain:


from langchain_community.tools import DuckDuckGoSearchRun


Create a function to initialize the search tool:


@st.cache_resource

def load_search_tool():

    """

    Initialize the DuckDuckGo search tool.

    Cached to avoid recreating on every interaction.

    """

    search = DuckDuckGoSearchRun()

    return search


Load the search tool:


# Initialize search tool

search_tool = load_search_tool()


Now create a function that performs searches and formats the results:


def search_internet(query):

    """

    Perform an internet search and return formatted results.

    """

    try:

        # Perform the search

        results = search_tool.run(query)

        

        # Format the results nicely

        formatted_results = f"Search results for '{query}':\n\n{results}"

        

        return formatted_results

        

    except Exception as e:

        return f"Search failed: {str(e)}"


This function tries to perform a search and handles any errors that might occur. The search_tool.run method returns a string containing search results, which we format and return.

Now we need to modify our conversation manager to use search results when available:


class ConversationManager:

    """

    Manages conversation history and generates responses.

    """

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages, search_results=None):

        """

        Format the conversation history into a prompt for the model.

        Optionally includes search results if provided.

        """

        # Take only the last max_history messages

        recent_messages = messages[-self.max_history:]

        

        # Build the conversation string

        conversation = ""

        

        # Add search results if provided

        if search_results:

            conversation += f"Reference Information:\n{search_results}\n\n"

        

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        conversation += "Assistant:"

        

        return conversation

    

    def generate_response(self, messages, search_results=None):

        """

        Generate a response based on conversation history.

        Optionally uses search results if provided.

        """

        # Format the conversation

        prompt = self.format_conversation(messages, search_results)

        

        # Generate response

        outputs = self.generator(

            prompt,

            max_new_tokens=150,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        # Extract the generated text

        generated_text = outputs[0]['generated_text']

        

        # Remove the prompt part

        response = generated_text[len(prompt):].strip()

        

        # Clean up

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


We modified the format_conversation and generate_response methods to accept optional search_results. When search results are provided, they are included in the prompt, giving the model additional context to generate better responses.


STEP SIX: BUILDING THE SIDEBAR WITH CONTROLS

Now let us add a sidebar to our interface with controls for search mode and session management. The sidebar will appear on the left side of the screen and contain buttons and toggles.

Add this code to create the sidebar:


# Sidebar for controls

with st.sidebar:

    st.header("⚙️ Settings")

    

    # Search toggle

    search_enabled = st.toggle(

        "🔍 Enable Internet Search",

        value=st.session_state.search_enabled,

        help="When enabled, the chatbot will search the internet for current information"

    )

    st.session_state.search_enabled = search_enabled

    

    # Display current status

    if search_enabled:

        st.success("Search Mode: ON")

        st.info("The chatbot will search the internet before responding")

    else:

        st.info("Search Mode: OFF")

        st.info("The chatbot will use only its training data")

    

    st.divider()

    

    # Session information

    st.header("📝 Session Info")

    st.text(f"Session ID: {st.session_state.session_id}")

    st.text(f"Messages: {len(st.session_state.messages)}")

    

    # Button to start new session

    if st.button("🔄 New Session"):

        save_conversation()

        st.session_state.messages = []

        st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")

        st.rerun()

    

    st.divider()

    

    # Button to save conversation

    if st.button("💾 Save Conversation"):

        save_conversation()

        st.success("Conversation saved!")

    

    # Button to load previous conversations

    st.header("📚 Previous Sessions")

    conversation_files = get_saved_conversations()

    

    if conversation_files:

        selected_file = st.selectbox(

            "Load a previous conversation:",

            conversation_files

        )

        

        if st.button("📂 Load Selected"):

            load_conversation(selected_file)

            st.rerun()

    else:

        st.text("No saved conversations yet")


This sidebar includes a toggle for search mode, displays session information, and provides buttons for managing conversations. The toggle lets users turn internet search on or off. The session info shows the unique ID and message count. Users can start a new session, which saves the current one and clears the screen. They can also manually save conversations or load previous ones.


STEP SEVEN: IMPLEMENTING SESSION PERSISTENCE

To save and load conversations, we need functions that write to and read from files. Let us create a directory for storing conversations and implement the necessary functions:


# Create directory for saved conversations if it doesn't exist

SAVE_DIR = "saved_conversations"

if not os.path.exists(SAVE_DIR):

    os.makedirs(SAVE_DIR)


Add these functions:


def save_conversation():

    """

    Save the current conversation to a JSON file.

    """

    if len(st.session_state.messages) == 0:

        return

    

    filename = f"{st.session_state.session_id}.json"

    filepath = os.path.join(SAVE_DIR, filename)

    

    conversation_data = {

        "session_id": st.session_state.session_id,

        "timestamp": datetime.now().isoformat(),

        "messages": st.session_state.messages

    }

    

    with open(filepath, 'w') as f:

        json.dump(conversation_data, f, indent=2)


def get_saved_conversations():

    """

    Get a list of all saved conversation files.

    """

    if not os.path.exists(SAVE_DIR):

        return []

    

    files = [f for f in os.listdir(SAVE_DIR) if f.endswith('.json')]

    # Sort by modification time, most recent first

    files.sort(key=lambda x: os.path.getmtime(os.path.join(SAVE_DIR, x)), reverse=True)

    return files


def load_conversation(filename):

    """

    Load a conversation from a saved file.

    """

    filepath = os.path.join(SAVE_DIR, filename)

    

    with open(filepath, 'r') as f:

        conversation_data = json.load(f)

    

    st.session_state.messages = conversation_data["messages"]

    st.session_state.session_id = conversation_data["session_id"]


These functions handle saving conversations as JSON files and loading them back. Each conversation is saved with its session ID as the filename. The JSON file contains the session ID, timestamp, and all messages. When loading, we restore the messages and session ID to session state.


STEP EIGHT: PUTTING IT ALL TOGETHER

Now we need to modify our chat input section to use all the components we have built. Replace the simple echo code with this complete implementation:


# Chat input

if prompt := st.chat_input("Type your message here..."):

    # Add user message to chat

    st.session_state.messages.append({"role": "user", "content": prompt})

    with st.chat_message("user"):

        st.markdown(prompt)

    

    # Generate response

    with st.chat_message("assistant"):

        with st.spinner("Thinking..."):

            # Check if search is enabled

            if st.session_state.search_enabled:

                # Perform internet search

                with st.status("Searching the internet...", expanded=True):

                    st.write("🔍 Performing web search...")

                    search_results = search_internet(prompt)

                    st.write("✅ Search complete!")

                

                # Generate response with search results

                response = conversation_manager.generate_response(

                    st.session_state.messages,

                    search_results=search_results

                )

            else:

                # Generate response without search

                response = conversation_manager.generate_response(

                    st.session_state.messages

                )

            

            # Display the response

            st.markdown(response)

    

    # Add assistant response to chat history

    st.session_state.messages.append({"role": "assistant", "content": response})

    

    # Auto-save after each interaction

    save_conversation()


This is where everything comes together. When a user sends a message, we first add it to our message history and display it. Then we check if search mode is enabled. If it is, we perform a web search and show the progress with status messages. We then generate a response using either just the conversation history or the conversation history plus search results. Finally, we display the response, add it to our history, and automatically save the conversation.


STEP NINE: ADDING ERROR HANDLING AND POLISH

Let us add some error handling and polish to make our chatbot more robust. Wrap the main chat functionality in a try-except block:


# Main chat interface

try:

    # Chat input

    if prompt := st.chat_input("Type your message here..."):

        # Add user message

        st.session_state.messages.append({"role": "user", "content": prompt})

        with st.chat_message("user"):

            st.markdown(prompt)

        

        # Generate response

        with st.chat_message("assistant"):

            try:

                with st.spinner("Thinking..."):

                    if st.session_state.search_enabled:

                        with st.status("Searching the internet...", expanded=True):

                            st.write("🔍 Performing web search...")

                            search_results = search_internet(prompt)

                            st.write("✅ Search complete!")

                        

                        response = conversation_manager.generate_response(

                            st.session_state.messages,

                            search_results=search_results

                        )

                    else:

                        response = conversation_manager.generate_response(

                            st.session_state.messages

                        )

                    

                    st.markdown(response)

            

            except Exception as e:

                error_msg = f"Sorry, I encountered an error: {str(e)}"

                st.error(error_msg)

                response = error_msg

        

        # Add response to history

        st.session_state.messages.append({"role": "assistant", "content": response})

        

        # Auto-save

        save_conversation()


except Exception as e:

    st.error(f"An error occurred: {str(e)}")


This error handling ensures that if something goes wrong, the app will not crash. Instead, it will show an error message to the user.

Let us also add a footer with helpful information:


# Footer

st.divider()

st.markdown("""

### 💡 Tips for Using This Chatbot:

- Toggle **Internet Search** ON to get current information from the web

- The chatbot remembers your conversation history during each session

- Click **New Session** to start fresh (your current conversation will be saved)

- Save important conversations using the **Save** button

- Load previous conversations from the sidebar

""")


THE COMPLETE WORKING CODE

Here is the complete code for our chatbot application. Save this as “chatbot_app.py”:


import streamlit as st

import torch

from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline

from langchain_community.tools import DuckDuckGoSearchRun

from datetime import datetime

import json

import os


# Page configuration

st.set_page_config(

    page_title="My AI Chatbot",

    page_icon="🤖",

    layout="wide"

)


# Create directory for saved conversations

SAVE_DIR = "saved_conversations"

if not os.path.exists(SAVE_DIR):

    os.makedirs(SAVE_DIR)


# Device detection function

def get_device():

    """Detect the best available device for running the model."""

    if torch.cuda.is_available():

        device = torch.device("cuda")

        device_name = torch.cuda.get_device_name(0)

        print(f"Using CUDA device: {device_name}")

        return device

    elif torch.backends.mps.is_available():

        device = torch.device("mps")

        print("Using Apple MPS device")

        return device

    else:

        device = torch.device("cpu")

        print("Using CPU device")

        return device


# Get device

device = get_device()


# Load model function

@st.cache_resource

def load_model():

    """Load the language model and tokenizer."""

    model_name = "gpt2"

    

    tokenizer = AutoTokenizer.from_pretrained(model_name)

    model = AutoModelForCausalLM.from_pretrained(

        model_name,

        torch_dtype=torch.float16 if device.type != "cpu" else torch.float32

    )

    model = model.to(device)

    

    generator = pipeline(

        "text-generation",

        model=model,

        tokenizer=tokenizer,

        device=device.type if device.type != "mps" else -1

    )

    

    return generator, tokenizer


# Load search tool

@st.cache_resource

def load_search_tool():

    """Initialize the DuckDuckGo search tool."""

    search = DuckDuckGoSearchRun()

    return search


# Conversation Manager class

class ConversationManager:

    """Manages conversation history and generates responses."""

    

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages, search_results=None):

        """Format conversation history into a prompt."""

        recent_messages = messages[-self.max_history:]

        conversation = ""

        

        if search_results:

            conversation += f"Reference Information:\n{search_results}\n\n"

        

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        conversation += "Assistant:"

        return conversation

    

    def generate_response(self, messages, search_results=None):

        """Generate a response based on conversation history."""

        prompt = self.format_conversation(messages, search_results)

        

        outputs = self.generator(

            prompt,

            max_new_tokens=150,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        generated_text = outputs[0]['generated_text']

        response = generated_text[len(prompt):].strip()

        

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


# Search function

def search_internet(query):

    """Perform an internet search and return results."""

    try:

        results = search_tool.run(query)

        formatted_results = f"Search results for '{query}':\n\n{results}"

        return formatted_results

    except Exception as e:

        return f"Search failed: {str(e)}"


# Conversation saving functions

def save_conversation():

    """Save the current conversation to a JSON file."""

    if len(st.session_state.messages) == 0:

        return

    

    filename = f"{st.session_state.session_id}.json"

    filepath = os.path.join(SAVE_DIR, filename)

    

    conversation_data = {

        "session_id": st.session_state.session_id,

        "timestamp": datetime.now().isoformat(),

        "messages": st.session_state.messages

    }

    

    with open(filepath, 'w') as f:

        json.dump(conversation_data, f, indent=2)


def get_saved_conversations():

    """Get list of all saved conversation files."""

    if not os.path.exists(SAVE_DIR):

        return []

    

    files = [f for f in os.listdir(SAVE_DIR) if f.endswith('.json')]

    files.sort(key=lambda x: os.path.getmtime(os.path.join(SAVE_DIR, x)), reverse=True)

    return files


def load_conversation(filename):

    """Load a conversation from a saved file."""

    filepath = os.path.join(SAVE_DIR, filename)

    

    with open(filepath, 'r') as f:

        conversation_data = json.load(f)

    

    st.session_state.messages = conversation_data["messages"]

    st.session_state.session_id = conversation_data["session_id"]


# Load model and tools

with st.spinner("Loading AI model... This may take a minute..."):

    generator, tokenizer = load_model()

    search_tool = load_search_tool()


# Initialize conversation manager

conversation_manager = ConversationManager(generator, tokenizer)


# Initialize session state

if 'messages' not in st.session_state:

    st.session_state.messages = []


if 'search_enabled' not in st.session_state:

    st.session_state.search_enabled = False


if 'session_id' not in st.session_state:

    st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")


# Title

st.title("🤖 My AI Chatbot")

st.markdown("Chat with an AI powered by Large Language Models")


# Sidebar

with st.sidebar:

    st.header("⚙️ Settings")

    

    search_enabled = st.toggle(

        "🔍 Enable Internet Search",

        value=st.session_state.search_enabled,

        help="When enabled, the chatbot will search the internet for current information"

    )

    st.session_state.search_enabled = search_enabled

    

    if search_enabled:

        st.success("Search Mode: ON")

        st.info("The chatbot will search the internet before responding")

    else:

        st.info("Search Mode: OFF")

        st.info("The chatbot will use only its training data")

    

    st.divider()

    

    st.header("📝 Session Info")

    st.text(f"Session ID: {st.session_state.session_id}")

    st.text(f"Messages: {len(st.session_state.messages)}")

    

    if st.button("🔄 New Session"):

        save_conversation()

        st.session_state.messages = []

        st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")

        st.rerun()

    

    st.divider()

    

    if st.button("💾 Save Conversation"):

        save_conversation()

        st.success("Conversation saved!")

    

    st.header("📚 Previous Sessions")

    conversation_files = get_saved_conversations()

    

    if conversation_files:

        selected_file = st.selectbox(

            "Load a previous conversation:",

            conversation_files

        )

        

        if st.button("📂 Load Selected"):

            load_conversation(selected_file)

            st.rerun()

    else:

        st.text("No saved conversations yet")


# Display chat messages

for message in st.session_state.messages:

    with st.chat_message(message["role"]):

        st.markdown(message["content"])


# Chat input

try:

    if prompt := st.chat_input("Type your message here..."):

        st.session_state.messages.append({"role": "user", "content": prompt})

        with st.chat_message("user"):

            st.markdown(prompt)

        

        with st.chat_message("assistant"):

            try:

                with st.spinner("Thinking..."):

                    if st.session_state.search_enabled:

                        with st.status("Searching the internet...", expanded=True):

                            st.write("🔍 Performing web search...")

                            search_results = search_internet(prompt)

                            st.write("✅ Search complete!")

                        

                        response = conversation_manager.generate_response(

                            st.session_state.messages,

                            search_results=search_results

                        )

                    else:

                        response = conversation_manager.generate_response(

                            st.session_state.messages

                        )

                    

                    st.markdown(response)

            

            except Exception as e:

                error_msg = f"Sorry, I encountered an error: {str(e)}"

                st.error(error_msg)

                response = error_msg

        

        st.session_state.messages.append({"role": "assistant", "content": response})

        save_conversation()


except Exception as e:

    st.error(f"An error occurred: {str(e)}")


# Footer

st.divider()

st.markdown("""

### 💡 Tips for Using This Chatbot:

- Toggle **Internet Search** ON to get current information from the web

- The chatbot remembers your conversation history during each session

- Click **New Session** to start fresh (your current conversation will be saved)

- Save important conversations using the **Save** button

- Load previous conversations from the sidebar

""")


RUNNING YOUR CHATBOT

Now that you have created your complete chatbot application, it is time to run it and see it in action. Make sure your virtual environment is still activated. In your terminal, navigate to the directory containing your “chatbot_app.py” file. Then run this command:


streamlit run chatbot_app.py


After a few seconds, your default web browser should automatically open and display your chatbot interface. If it does not open automatically, look at your terminal output. Streamlit will print a local URL like “http://localhost:8501”. You can copy and paste this URL into your browser.

The first time you run the chatbot, it will take a minute or two to load the AI model. You will see a loading message while this happens. Once the model is loaded, you are ready to start chatting.

Try having a conversation with your chatbot. Type a message in the input box at the bottom of the screen and press Enter. The chatbot will think for a moment and then respond. Notice how it remembers previous messages in the conversation - you can refer back to things you discussed earlier and it will understand the context.

Now try enabling internet search by clicking the toggle in the sidebar. Ask a question about something current or recent, like “What are the latest developments in space exploration?” The chatbot will perform a web search, and you will see status messages showing this process. The response should incorporate information from the search results.

Experiment with starting new sessions and loading previous conversations. Each session gets saved automatically, so you never lose your chat history. You can come back days later and resume old conversations.


UNDERSTANDING HOW THE COMPONENTS WORK TOGETHER

Now that you have a working chatbot, let us review how all the pieces fit together to create the complete system. Understanding this architecture will help you customize and extend your chatbot.

The application starts by loading the AI model and search tool. This happens once when the app first starts, thanks to the Streamlit caching decorators. The model is moved to the best available device - CUDA for NVIDIA GPUs, MPS for Apple Silicon, or CPU as a fallback.

When you type a message, Streamlit captures it and adds it to the session state messages list. This list persists across interactions because of how Streamlit’s session state works. Your message is immediately displayed in the chat interface.

Next, the application checks whether internet search is enabled. This is controlled by the toggle button in the sidebar, which updates a session state variable. If search is on, the DuckDuckGo search tool is called with your message as the query. The search results are retrieved and formatted into a string.

The conversation manager then takes over. It looks at the message history and optionally the search results, and formats them into a prompt string that the AI model can understand. This prompt includes recent messages to give context, and the search results if available to provide additional information.

The formatted prompt is sent to the AI model through the text generation pipeline. The model processes this input and generates a response. The generation parameters control aspects like response length and creativity. The model’s raw output includes both the input prompt and the generated response, so we extract just the new text.

This generated response is displayed in the chat interface and added to the message history. Finally, the entire conversation is automatically saved to a JSON file in the saved_conversations directory. This happens after every interaction, so you never lose data even if you close the browser.

The sidebar provides controls that modify this flow. The search toggle changes whether internet lookup happens. The new session button saves the current conversation and resets the message history. The load conversation feature retrieves a previously saved conversation and restores it to the current session.


TESTING YOUR CHATBOT THOROUGHLY

Before you share your chatbot or rely on it for important tasks, you should test it thoroughly to understand its capabilities and limitations. Here are some testing scenarios you should try.

First, test basic conversational abilities. Have a multi-turn conversation where you ask follow-up questions that refer to previous messages. For example, ask “What is Python?”, then “What are its main uses?”, then “Which one would you recommend for a beginner?” The chatbot should maintain context throughout.

Second, test the search functionality. Turn on internet search and ask questions about recent events or current information that would not be in the model’s training data. Compare the responses with search enabled versus disabled to see the difference. Good test questions include “What is the current date?”, “What are today’s top news headlines?”, or “What is the weather in [your city]?”

Third, test the session management. Create a conversation, start a new session, then load the previous conversation to verify it was saved correctly. Try loading multiple different conversations to ensure the system correctly switches between them.

Fourth, test error handling. Try to break the system by entering very long messages, special characters, or unusual inputs. The error handling should catch problems and display user-friendly messages rather than crashing.

Fifth, test hardware acceleration. Look at the console output when the app starts to verify it detected your hardware correctly. If you have a GPU, the model should run faster than on CPU.

Make notes about what works well and what could be improved. Every limitation you discover is an opportunity for enhancement.


TROUBLESHOOTING COMMON ISSUES


As you work with your chatbot, you might encounter some common issues. Here are solutions to the most frequent problems.

If the model takes a very long time to load, this is normal the first time. The Hugging Face Transformers library downloads the model files from the internet, which can take several minutes depending on your connection speed. Subsequent runs will be much faster because the model is cached locally.

If you get an error about not having enough memory, the AI model might be too large for your system. Try using a smaller model by changing the model_name variable to “distilgpt2”, which is a compressed version of GPT-2 that uses less memory. Alternatively, make sure you close other memory-intensive applications.

If the search functionality does not work, verify that you installed the duckduckgo-search package correctly. You can test it separately by opening a Python console and running “from langchain_community.tools import DuckDuckGoSearchRun” followed by “search = DuckDuckGoSearchRun()” and “search.run(‘test query’)”. If this works, the problem is elsewhere in your code.

If saved conversations do not load, check that the saved_conversations directory exists and contains JSON files. You can open these files in a text editor to verify they are formatted correctly. The load_conversation function expects specific keys in the JSON structure.

If responses seem inconsistent or low quality, try adjusting the generation parameters. Lowering the temperature makes responses more focused and predictable. Increasing max_new_tokens allows longer responses. Experiment with different values to find what works best.

If the application crashes with errors about device placement, check that your PyTorch installation matches your hardware. For NVIDIA GPUs, ensure you installed the CUDA-enabled version. For Apple Silicon, verify you have the latest PyTorch version with MPS support.


FUTURE ENHANCEMENTS AND EXTENSIONS

Now that you have a working chatbot, there are many ways you can enhance and extend it. Here are some ideas for improvements you might want to implement.

One significant enhancement would be using a more powerful language model. GPT-2 is good for learning, but larger models produce better responses. You could upgrade to GPT-2 Medium or Large, or try other models like Microsoft’s DialoGPT which is specifically trained for conversations. More advanced users might experiment with models like FLAN-T5 or LLaMA.

You could add support for multiple search engines. Currently we only use DuckDuckGo, but LangChain supports other search tools like Google Search, Wikipedia, and Wolfram Alpha. You could let users choose which search engine to use, or have the chatbot automatically select the best one based on the question type.

Another useful feature would be document upload. You could let users upload PDF or text files, and the chatbot could answer questions about the content. This would require adding document loading and text splitting capabilities, which LangChain provides through its document loaders and text splitters.

You could implement a more sophisticated conversation history system. Instead of just keeping recent messages, you could use embeddings to find the most relevant past messages, or implement summarization to condense long conversations. This would let the chatbot maintain context over much longer conversations without exceeding token limits.

Adding user authentication would let multiple people use the same chatbot instance while keeping their conversations separate and private. You could use Streamlit’s authentication components or integrate with existing authentication systems.

You could add voice input and output. The speech_recognition library can convert speech to text, and text-to-speech libraries like pyttsx3 or gTTS can convert responses to audio. This would create a more natural, hands-free interaction model.

Another enhancement would be adding the ability to generate images using models like Stable Diffusion when users ask for visual content. You could integrate this as another tool the chatbot can use when appropriate.

You could also implement a rating system where users can rate responses, and use this feedback to fine-tune the model or select better generation parameters. This creates a feedback loop for continuous improvement.

For more advanced users, you could implement Retrieval Augmented Generation or RAG, which combines the language model with a vector database of documents. This allows the chatbot to draw from a large knowledge base while still generating natural language responses.

You might add support for different languages by using multilingual models or translation APIs. This would let users chat in their preferred language.

Finally, you could deploy your chatbot to the cloud so others can access it. Streamlit offers Streamlit Cloud for easy deployment, or you could use platforms like Heroku, AWS, or Google Cloud Platform.


BEST PRACTICES FOR LLM APPLICATION DEVELOPMENT

As you continue developing your chatbot and other AI applications, keep these best practices in mind to create high-quality, maintainable systems.

Always handle errors gracefully. AI models can be unpredictable, and external services like search APIs can fail. Wrap potentially problematic code in try-except blocks and provide helpful error messages to users. Never let exceptions crash your application.

Use caching strategically. Loading models and initializing components is expensive. The Streamlit cache_resource decorator ensures these operations happen only once, dramatically improving performance. However, be careful not to cache things that should change between interactions.

Monitor resource usage. AI models consume significant memory and processing power. Limit conversation history length to avoid exceeding memory limits. Consider implementing cleanup mechanisms for long-running applications.

Document your code thoroughly. Even though this tutorial includes extensive comments, in real projects you should add even more documentation. Explain why you made certain design decisions, not just what the code does. Future you will thank present you.

Test with diverse inputs. Users will interact with your chatbot in unexpected ways. Test edge cases, unusual inputs, and potential misuse scenarios. Ensure your application behaves appropriately in all situations.

Be transparent about limitations. AI models can produce incorrect or biased responses. Make sure users understand they are interacting with an AI and should verify important information. Consider adding disclaimers or warnings for sensitive use cases.

Respect privacy and security. If your chatbot handles personal information, ensure you follow data protection best practices. Do not log sensitive information, use encryption for data at rest and in transit, and provide clear privacy policies.

Optimize for user experience. Response time matters. If generation takes too long, users will lose interest. Show progress indicators during long operations. Make the interface intuitive and provide helpful guidance.

Keep models and libraries updated. The AI field evolves rapidly. Regularly update your dependencies to get bug fixes, performance improvements, and new features. However, test thoroughly after updates to ensure nothing breaks.

Learn from user interactions. If possible, collect anonymized usage data to understand how people use your chatbot. This helps identify which features are valuable and where users struggle.


UNDERSTANDING THE ETHICAL IMPLICATIONS

As you develop AI applications, it is important to understand the ethical implications and responsibilities that come with this technology. AI systems can have significant impacts on people’s lives, and developers have a responsibility to create systems that are beneficial and minimize harm.

Large Language Models can sometimes generate biased, incorrect, or harmful content. This happens because they are trained on internet text, which contains human biases and errors. As a developer, you should be aware of these limitations and implement safeguards. Consider adding content filtering to block inappropriate outputs, and be transparent with users about potential issues.

Privacy is a critical concern. Our chatbot saves conversation history, which might contain sensitive information. Make sure users understand what data is collected and how it is used. Implement appropriate security measures to protect this data. Give users control over their data, including the ability to delete their information.

The environmental impact of AI is another consideration. Training large models requires enormous computational resources and energy. While you are not training models in this tutorial, be mindful of energy efficiency in your applications. Using smaller models when possible and optimizing code can reduce environmental impact.

Accessibility is important too. Not everyone interacts with technology in the same way. Consider how users with disabilities might use your chatbot. Can it be used with screen readers? Can users adjust text size? Thinking about accessibility from the start makes your application more inclusive.

Be honest about what your chatbot can and cannot do. Do not overstate its capabilities or present it as something it is not. If it is a simple model with limited abilities, be clear about that. Users should understand they are interacting with an AI, not a human expert.

Think about potential misuse. Could your chatbot be used to spread misinformation, create spam, or harm others? While you cannot prevent all misuse, you can design your system to make harmful uses more difficult. This might include rate limiting, content filtering, or requiring authentication.


LEARNING RESOURCES FOR CONTINUED GROWTH

This tutorial has given you a solid foundation in building LLM applications, but there is much more to learn. Here are resources to help you continue your learning journey.

The official documentation for the libraries we used is essential reading. The Streamlit documentation at docs.streamlit.io provides detailed guides and examples for building web applications. The LangChain documentation at python.langchain.com covers all aspects of the framework, including advanced features we did not explore. The Hugging Face Transformers documentation at huggingface.co/docs/transformers is comprehensive and includes tutorials for many different tasks.

Online courses can provide structured learning paths. Coursera and Udemy offer courses on natural language processing, deep learning, and LLM applications. Many are taught by university professors and industry experts. Some are free, others require payment but often provide certificates.

YouTube has excellent tutorial channels. Channels like “Corey Schafer” and “Tech With Tim” have Python tutorials. For AI-specific content, “AssemblyAI” and “1littlecoder” create excellent tutorials on building AI applications. “Hugging Face” has their own channel with model demonstrations and technical talks.

Reading research papers, while challenging, helps you understand the latest developments. Papers like “Attention Is All You Need” which introduced Transformers, and GPT papers from OpenAI provide insights into how these models work. The ArXiv website hosts pre-publication papers on AI and machine learning.

Books provide in-depth knowledge. “Natural Language Processing with Transformers” by Lewis Tunstall and others covers modern NLP techniques. “Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” by Aurélien Géron is an excellent introduction to machine learning. “Deep Learning” by Ian Goodfellow is more advanced but comprehensive.

Online communities are valuable for getting help and sharing knowledge. The Streamlit forum, LangChain Discord server, and Hugging Face forums are active communities where you can ask questions and learn from others. Reddit communities like r/MachineLearning and r/LanguageTechnology discuss AI developments and applications.

GitHub is an incredible resource for learning from real projects. Search for repositories using the technologies we covered. Reading other people’s code is one of the best ways to learn new techniques and best practices. You can also contribute to open source projects to gain experience.

Experiment with different models and techniques. The best way to learn is by doing. Try modifying your chatbot to add new features. Experiment with different models from Hugging Face. Build completely new applications using what you have learned. Each project will teach you something new.


CONCLUSION

Congratulations on building your own AI chatbot. You have learned a tremendous amount in this tutorial, from understanding Large Language Models to implementing a complete application with a web interface, conversation management, internet search capabilities, and hardware acceleration.

The chatbot you built is not just a learning exercise - it is a functional application that demonstrates real AI capabilities. You have conversation history that maintains context across multiple exchanges. You have session management that saves and loads conversations. You have internet search that provides up-to-date information. And you have hardware detection that ensures your application runs efficiently on different systems.

More importantly, you have learned the fundamental concepts and techniques used in modern AI application development. You understand how to work with pre-trained models, how to build user interfaces with Streamlit, how to use frameworks like LangChain to simplify complex tasks, and how to integrate multiple components into a cohesive system.

The skills you have learned are applicable far beyond this specific chatbot. The patterns and techniques we used - component-based architecture, session management, error handling, and hardware optimization - apply to many types of applications. The experience of reading documentation, debugging problems, and iterating on features is essential to all software development.

As you continue your journey in AI and software development, remember that learning is a continuous process. Technology evolves rapidly, especially in AI. New models, frameworks, and techniques emerge constantly. Stay curious, keep experimenting, and never stop learning.

Your chatbot is just the beginning. With the foundation you have built, you can create increasingly sophisticated applications. Maybe you will build a chatbot that helps students learn difficult subjects. Or a virtual assistant that helps people manage their daily tasks. Or a research tool that helps scientists explore complex questions. The possibilities are limited only by your imagination and dedication.

Thank you for working through this tutorial. We hope you found it educational, engaging, and inspiring. Now go forth and build amazing things with AI.


(to be CONTINUED)


ADDENDUM: RESOURCES


BOOKS FOR FURTHER LEARNING

“Natural Language Processing with Transformers” by Lewis Tunstall, Leandro von Werra, and Thomas Wolf. Published by O’Reilly Media. This book is specifically focused on using Transformers for NLP tasks and covers practical applications. It is very hands-on and includes numerous code examples using the Hugging Face library.

“Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” by Aurélien Géron. Published by O’Reilly Media. While not specifically about LLMs, this book provides essential machine learning fundamentals that underpin AI applications. The second edition covers deep learning extensively.

“Deep Learning” by Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Published by MIT Press. This is a comprehensive textbook that covers the mathematical and theoretical foundations of deep learning. It is more academic than the others but provides deep understanding.

“Building Machine Learning Powered Applications” by Emmanuel Ameisen. Published by O’Reilly Media. This book focuses on the practical aspects of deploying machine learning models in production, including many considerations we touched on in this tutorial.

“Speech and Language Processing” by Dan Jurafsky and James H. Martin. Available free online. This is a comprehensive textbook on natural language processing that covers both traditional and modern approaches. The third edition includes coverage of neural approaches.


WEBSITES AND ONLINE DOCUMENTATION

Official Streamlit Documentation at https://docs.streamlit.io provides comprehensive guides, API references, and examples for building web applications with Streamlit.

LangChain Documentation at https://python.langchain.com/docs/get_started/introduction covers the entire LangChain framework with tutorials, how-to guides, and conceptual explanations.

Hugging Face Documentation at https://huggingface.co/docs provides documentation for all Hugging Face libraries including Transformers, Datasets, and Accelerate. The Hugging Face Hub at https://huggingface.co/models hosts thousands of pre-trained models.

PyTorch Documentation at https://pytorch.org/docs/stable/index.html covers the PyTorch deep learning framework in detail, including guides for using CUDA and MPS acceleration.

The official Python documentation at https://docs.python.org/3/ is an essential reference for Python itself.


ONLINE COURSES AND TUTORIALS

Coursera offers several relevant courses. “Natural Language Processing” by DeepLearning.AI covers modern NLP techniques. “Deep Learning Specialization” by Andrew Ng is a comprehensive introduction to deep learning. “Machine Learning” by Andrew Ng is a classic introduction to machine learning fundamentals.

Fast.ai offers free courses including “Practical Deep Learning for Coders” which takes a code-first approach to learning deep learning. The course includes a chapter on natural language processing.

The Hugging Face Course at https://huggingface.co/course/chapter1/1 is a free comprehensive course on using Transformers for various NLP tasks. It includes hands-on exercises and covers advanced topics.

YouTube has many excellent channels. “3Blue1Brown” explains mathematical concepts visually, which is helpful for understanding how neural networks work. “Sentdex” has practical Python and machine learning tutorials. “Two Minute Papers” summarizes recent AI research papers in accessible formats.


RESEARCH PAPERS AND TECHNICAL BLOGS

“Attention Is All You Need” by Vaswani et al. introduced the Transformer architecture that powers modern LLMs. Available on ArXiv at https://arxiv.org/abs/1706.03762

“Language Models are Few-Shot Learners” describes GPT-3 and demonstrates the capabilities of large language models. Available at https://arxiv.org/abs/2005.14165

The OpenAI Blog at https://openai.com/blog publishes accessible explanations of their research and models.

The Hugging Face Blog at https://huggingface.co/blog features technical tutorials and announcements about new models and techniques.

Google AI Blog at https://ai.googleblog.com discusses Google’s AI research including work on language models.

Anthropic’s Research page at https://www.anthropic.com/research publishes papers about AI safety and capabilities.


COMMUNITIES AND FORUMS

The Streamlit Forum at https://discuss.streamlit.io is an active community where you can ask questions about Streamlit and share your projects.

The LangChain GitHub Discussions at https://github.com/langchain-ai/langchain/discussions is where the LangChain community discusses the framework.

Hugging Face Forums at https://discuss.huggingface.co covers topics related to models, datasets, and the Transformers library.

Reddit has several relevant communities. r/MachineLearning discusses machine learning research and applications. r/LanguageTechnology focuses on NLP and computational linguistics. r/learnmachinelearning is for beginners asking questions and sharing resources.

Discord servers exist for many AI frameworks and tools. The LangChain Discord and Hugging Face Discord are particularly active.

Stack Overflow at https://stackoverflow.com is invaluable for getting help with specific programming problems.


TOOLS AND PLATFORMS

Google Colab at https://colab.research.google.com provides free access to GPUs for running machine learning code in Jupyter notebooks.

Kaggle at https://www.kaggle.com offers free GPU access, datasets, and competitions. It is a good place to practice machine learning.

GitHub at https://github.com is essential for version control and for exploring open source projects. Learning Git is valuable for any programmer.

Weights and Biases at https://wandb.ai provides tools for tracking machine learning experiments.

Streamlit Cloud at https://streamlit.io/cloud offers free deployment for Streamlit applications.