Saturday, October 10, 2026

BUILDING YOUR OWN LLM CHATBOT IN PYTHON - PART 1/2: FROM FIRST STEPS TO WORKING IMPLEMENTATION


Part 1/2

INTRODUCTION: WELCOME TO THE WORLD OF AI CHATBOTS

Imagine having a conversation with a computer program that understands you, remembers what you talked about earlier, and can even search the internet to give you up-to-date information. This is exactly what we are going to build together in this tutorial. We will create a fully functional chatbot powered by Large Language Models, or LLMs for short.

Large Language Models are artificial intelligence systems that have been trained on massive amounts of text from books, websites, and other sources. They can understand human language and generate responses that sound remarkably human-like. You have probably used chatbots like ChatGPT or Claude before. In this tutorial, you will learn how to build your own version from scratch.

What makes this tutorial special is that we will build a chatbot that has real, practical features. Our chatbot will remember previous conversations, so you can have natural back-and-forth discussions just like texting with a friend. It will save different chat sessions so you can come back to old conversations later. Most excitingly, it will have a special internet search mode that allows it to look up current information online when you activate it with an icon, giving you the best of both worlds - the knowledge in its training and fresh information from the web.

We will also make sure our chatbot works efficiently on different types of computers. Whether you have a Mac with Apple Silicon chips or a Windows or Linux computer with an NVIDIA graphics card, our code will automatically detect and use the best available hardware to run the AI model quickly.

This tutorial is designed for high school students who have some basic Python programming knowledge. You should be comfortable with variables, functions, and simple data structures like lists and dictionaries. Do not worry if you have never worked with AI before - we will explain everything step by step, and every piece of code will be fully explained so you understand exactly what it does and why.


UNDERSTANDING THE TECHNOLOGY STACK

Before we start coding, let us understand the main tools and frameworks we will use. Think of building software like constructing a house - you need different materials and tools for different parts. Our chatbot will use several specialized tools, each designed for a specific purpose.

The first major tool is Streamlit. Streamlit is a Python framework that makes it incredibly easy to create web applications. Normally, building a web interface requires knowledge of HTML, CSS, and JavaScript. With Streamlit, you can create beautiful web interfaces using only Python. This is perfect for us because we can focus on the chatbot logic without getting bogged down in web development details. Streamlit will handle creating the chat window, displaying messages, and managing user input.

Next, we have LangChain, which is a framework specifically designed for building applications with Large Language Models. LangChain provides pre-built components for common tasks like managing conversation history, calling different AI models, and integrating external tools like web search. Instead of writing hundreds of lines of code from scratch, we can use LangChain’s ready-made components and focus on making our chatbot unique and useful.

For the actual AI brain of our chatbot, we will use the Transformers library from Hugging Face. Hugging Face is a company that has made thousands of pre-trained AI models available to developers for free. The Transformers library makes it easy to download these models and use them in your own applications. We will use a relatively small model that can run on regular computers rather than requiring expensive cloud servers.

For internet search capabilities, we will use DuckDuckGo Search. DuckDuckGo is a privacy-focused search engine, and LangChain provides a tool that lets our chatbot perform web searches. When you activate search mode, the chatbot will query DuckDuckGo and use the search results to give you current information.

Finally, PyTorch is the deep learning framework that powers the AI model. PyTorch includes support for different types of hardware acceleration. On Apple computers with M{1,2,3,4,5,6} chips, it can use Metal Performance Shaders or MPS for short to run the model faster. On computers with NVIDIA graphics cards, it uses CUDA. Our code will automatically detect what hardware you have and use it optimally.


SETTING UP YOUR DEVELOPMENT ENVIRONMENT

The first step in any programming project is setting up your development environment. This means installing Python and all the necessary libraries. We will walk through this process step by step.

First, make sure you have Python installed on your computer. You need Python version three point eight or higher. You can check if Python is installed by opening your terminal or command prompt and typing “python –version”. If Python is not installed, download it from python.org and follow the installation instructions for your operating system.

Once Python is installed, we strongly recommend creating a virtual environment. A virtual environment is like a separate workspace for your project that keeps all your project’s libraries isolated from other Python projects on your computer. This prevents conflicts between different versions of libraries. To create a virtual environment, open your terminal, navigate to the folder where you want to keep your project, and run these commands. On Windows, type “python -m venv chatbot_env” to create the environment, then “chatbot_env\Scripts\activate” to activate it. On Mac or Linux, type “python3 -m venv chatbot_env” to create it, then “source chatbot_env/bin/activate” to activate it.

When your virtual environment is activated, you will see its name in parentheses at the start of your command prompt. This tells you that any packages you install will go into this isolated environment rather than affecting your system-wide Python installation.


INSTALLING REQUIRED PACKAGES

Now we need to install all the Python packages our chatbot will use. We will install these one by one so you understand what each package does.

First, install Streamlit by typing “pip install streamlit”. This installs the web framework we will use to build our user interface.

Next, install LangChain and its community packages with “pip install langchain langchain-community”. LangChain Community contains additional integrations including the DuckDuckGo search tool.

Now install the Transformers library from Hugging Face with “pip install transformers”. This gives us access to pre-trained language models.

Install PyTorch, which is the deep learning framework that runs the models. The installation command depends on your system. For Mac with Apple Silicon, type “pip install torch torchvision torchaudio”. For Windows or Linux with NVIDIA GPU, visit pytorch.org to get the specific command for your CUDA version. For CPU-only systems, type “pip install torch torchvision torchaudio –index-url https://download.pytorch.org/whl/cpu”.

Install the Accelerate library with “pip install accelerate”. This library helps optimize model loading and inference across different hardware.

Install the DuckDuckGo search tool with “pip install duckduckgo-search”. This enables web search functionality.

Finally, install the Sentence Transformers library with “pip install sentence-transformers”. This provides efficient text embedding models that we might use for future enhancements.

If you want to install all these packages at once, you can create a file called “requirements.txt” with all the package names, then run “pip install -r requirements.txt”. Here is what your requirements.txt file should contain:


streamlit==1.29.0

langchain==0.1.0

langchain-community==0.0.13

transformers==4.36.0

torch==2.1.0

accelerate==0.25.0

duckduckgo-search==4.1.0

sentence-transformers==2.2.2


UNDERSTANDING HOW OUR CHATBOT WILL WORK

Before we start writing code, let us understand the architecture of our chatbot. Architecture means the overall structure and design of how different parts work together.

Our chatbot will have three main layers. The first layer is the user interface layer, which is what you see and interact with. This includes the chat window where messages appear, the text input box where you type your questions, and the search icon that toggles internet search mode. Streamlit handles all of this for us.

The second layer is the logic layer, which contains the brain of our application. This layer manages conversation history, decides when to use internet search, and coordinates between different components. This is where most of our custom code will live.

The third layer is the model layer, which contains the actual AI model that generates responses. This layer also includes the search tool for when internet lookup is needed. LangChain and Transformers libraries handle this layer.

Our chatbot will work like this. When you type a message and press enter, the user interface layer captures your input. The logic layer checks if search mode is activated. If search mode is off, the chatbot uses only the AI model to generate a response based on its training data and the conversation history. If search mode is on, the logic layer first performs a web search using DuckDuckGo, then provides both your question and the search results to the AI model, which formulates a comprehensive answer.

The conversation history is stored in memory during your session. Each session gets a unique identifier, and we save the history to a file when you end the session. This way, you can resume conversations later.


STEP ONE: CREATING THE BASIC STREAMLIT INTERFACE

Let us start by creating a simple Streamlit interface. This will be the foundation that we build upon. Create a new file called “chatbot_app.py” and open it in your text editor.

At the top of the file, we import the necessary libraries:


import streamlit as st

import torch

from datetime import datetime

import json

import os


Now let us configure the Streamlit page. This sets the title that appears in your browser tab and the layout:


st.set_page_config(

    page_title="My AI Chatbot",

    page_icon="🤖",

    layout="wide"

)


Next, we add a title to our page:


st.title("🤖 My AI Chatbot")

st.markdown("Chat with an AI powered by Large Language Models")


Streamlit uses something called session state to remember information between interactions. When a user clicks a button or types something, Streamlit reruns your entire script. Session state lets us preserve data across these reruns. We will initialize our session state variables:


# Initialize session state variables

if 'messages' not in st.session_state:

    st.session_state.messages = []


if 'search_enabled' not in st.session_state:

    st.session_state.search_enabled = False


if 'session_id' not in st.session_state:

    st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")


The messages list will store all the conversation messages. The search_enabled flag tells us whether internet search is currently activated. The session_id is a unique identifier for this conversation session based on the current date and time.

Now let us create the chat interface. Streamlit provides a special container for displaying chat messages:


# Display chat messages

for message in st.session_state.messages:

    with st.chat_message(message["role"]):

        st.markdown(message["content"])


This loop goes through all our saved messages and displays them. Each message has a role, which is either “user” for messages you typed or “assistant” for responses from the chatbot. The chat_message function displays the message with an appropriate icon.

Finally, we add an input box where users can type their messages:


# Chat input

if prompt := st.chat_input("Type your message here..."):

    # Add user message to chat

    st.session_state.messages.append({"role": "user", "content": prompt})

    with st.chat_message("user"):

        st.markdown(prompt)

    

    # For now, just echo the message back

    response = f"You said: {prompt}"

    st.session_state.messages.append({"role": "assistant", "content": response})

    with st.chat_message("assistant"):

        st.markdown(response)


The walrus operator “:=” assigns the user’s input to the prompt variable and checks if it is not empty, all in one line. If the user typed something, we add it to our messages list, display it, create a simple response, and display that too.

You can test this basic interface by saving the file and running “streamlit run chatbot_app.py” in your terminal. A web browser should open showing your chatbot interface. It will not be very smart yet - it just echoes back what you type - but you can see the chat interface working.


STEP TWO: ADDING HARDWARE DETECTION FOR GPU SUPPORT

Before we integrate the AI model, we need to write code that automatically detects what hardware acceleration is available. This is important because it ensures our chatbot runs as fast as possible on any computer.

Add this function near the top of your file, after the imports:


def get_device():

    """

    Detect the best available device for running the model.

    Returns a torch device object.

    """

    # Check for CUDA (NVIDIA GPU)

    if torch.cuda.is_available():

        device = torch.device("cuda")

        device_name = torch.cuda.get_device_name(0)

        print(f"Using CUDA device: {device_name}")

        return device

    

    # Check for MPS (Apple Silicon)

    elif torch.backends.mps.is_available():

        device = torch.device("mps")

        print("Using Apple MPS device")

        return device

    

    # Fall back to CPU

    else:

        device = torch.device("cpu")

        print("Using CPU device")

        return device


This function tries to detect GPU acceleration in order of preference. NVIDIA GPUs with CUDA are generally the fastest for AI models, so we check for that first. If we are on a Mac with Apple Silicon, we check for MPS availability. If neither is available, we fall back to using the CPU, which works on any computer but is slower.

Let us call this function and store the device:


# Get the best available device

device = get_device()


Now our chatbot knows what hardware to use for running the AI model.


STEP THREE: LOADING THE LANGUAGE MODEL

Now comes the exciting part - loading an actual AI language model. We will use a model called GPT-2, which is a smaller model that can run on regular computers. While it is not as powerful as models like GPT-4, it is perfect for learning and can run locally without needing expensive cloud services.

First, let us import the necessary classes from the Transformers library:


from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline


Now we will create a function that loads the model with caching so it only loads once:


@st.cache_resource

def load_model():

    """

    Load the language model and tokenizer.

    Uses Streamlit caching to avoid reloading on every interaction.

    """

    model_name = "gpt2"  # Using GPT-2 as our base model

    

    # Load the tokenizer

    # The tokenizer converts text into numbers the model can understand

    tokenizer = AutoTokenizer.from_pretrained(model_name)

    

    # Load the model

    model = AutoModelForCausalLM.from_pretrained(

        model_name,

        torch_dtype=torch.float16 if device.type != "cpu" else torch.float32

    )

    

    # Move model to the detected device (GPU or CPU)

    model = model.to(device)

    

    # Create a text generation pipeline

    generator = pipeline(

        "text-generation",

        model=model,

        tokenizer=tokenizer,

        device=device.type if device.type != "mps" else -1

    )

    

    return generator, tokenizer


The @st.cache_resource decorator is important. It tells Streamlit to run this function only once and remember the result. Without this, the model would reload every time you send a message, which would be extremely slow.

Let us break down what this function does. First, it specifies the model name. GPT-2 is available in several sizes, and we are using the smallest version. The tokenizer is loaded first - it converts human text into numbers that the model understands. Then we load the actual model.

The torch_dtype parameter specifies the precision of numbers used in the model. Float16 uses half-precision math, which is faster on GPUs and uses less memory. We only use this on GPU though - CPUs work better with float32.

We move the model to our detected device using the to method. Then we create a pipeline, which is a high-level interface that handles all the details of generating text.

Now let us actually load the model:


# Load the model when the app starts

with st.spinner("Loading AI model... This may take a minute..."):

    generator, tokenizer = load_model()


st.success("Model loaded successfully!")


The spinner shows a loading message while the model loads. This is important because loading can take 30 seconds to a minute the first time. The success message confirms everything is ready.


STEP FOUR: CREATING THE CONVERSATION MANAGER

To have natural conversations, we need to manage the conversation history properly. The AI model needs context from previous messages to generate relevant responses. Let us create a class that handles this:


class ConversationManager:

    """

    Manages conversation history and generates responses.

    """

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages):

        """

        Format the conversation history into a prompt for the model.

        """

        # Take only the last max_history messages to avoid exceeding token limits

        recent_messages = messages[-self.max_history:]

        

        # Build the conversation string

        conversation = ""

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        # Add the prompt for the next assistant response

        conversation += "Assistant:"

        

        return conversation

    

    def generate_response(self, messages):

        """

        Generate a response based on conversation history.

        """

        # Format the conversation

        prompt = self.format_conversation(messages)

        

        # Generate response

        outputs = self.generator(

            prompt,

            max_new_tokens=100,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        # Extract the generated text

        generated_text = outputs[0]['generated_text']

        

        # Remove the prompt part to get only the new response

        response = generated_text[len(prompt):].strip()

        

        # Clean up the response

        # Sometimes the model generates multiple lines, we want only the first response

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


This class encapsulates all the logic for managing conversations and generating responses. The format_conversation method creates a prompt string that includes the recent conversation history. We limit this to the last few messages to avoid exceeding the model’s maximum input length.

The generate_response method calls the model to generate a response. Let us understand the parameters. The max_new_tokens parameter limits how long the response can be. The temperature parameter controls randomness - lower values make responses more focused and deterministic, higher values make them more creative. We set do_sample to True to enable random sampling, which makes responses more natural. The pad_token_id parameter handles technical details about how sequences are padded.

After generation, we extract just the new text that the model added and clean it up.

Let us create an instance of our conversation manager:


# Initialize conversation manager

conversation_manager = ConversationManager(generator, tokenizer)


STEP FIVE: ADDING INTERNET SEARCH CAPABILITY

Now let us add the ability to search the internet. This is what makes our chatbot really useful - it can provide up-to-date information instead of being limited to its training data.

First, import the search tool from LangChain:


from langchain_community.tools import DuckDuckGoSearchRun


Create a function to initialize the search tool:


@st.cache_resource

def load_search_tool():

    """

    Initialize the DuckDuckGo search tool.

    Cached to avoid recreating on every interaction.

    """

    search = DuckDuckGoSearchRun()

    return search


Load the search tool:


# Initialize search tool

search_tool = load_search_tool()


Now create a function that performs searches and formats the results:


def search_internet(query):

    """

    Perform an internet search and return formatted results.

    """

    try:

        # Perform the search

        results = search_tool.run(query)

        

        # Format the results nicely

        formatted_results = f"Search results for '{query}':\n\n{results}"

        

        return formatted_results

        

    except Exception as e:

        return f"Search failed: {str(e)}"


This function tries to perform a search and handles any errors that might occur. The search_tool.run method returns a string containing search results, which we format and return.

Now we need to modify our conversation manager to use search results when available:


class ConversationManager:

    """

    Manages conversation history and generates responses.

    """

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages, search_results=None):

        """

        Format the conversation history into a prompt for the model.

        Optionally includes search results if provided.

        """

        # Take only the last max_history messages

        recent_messages = messages[-self.max_history:]

        

        # Build the conversation string

        conversation = ""

        

        # Add search results if provided

        if search_results:

            conversation += f"Reference Information:\n{search_results}\n\n"

        

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        conversation += "Assistant:"

        

        return conversation

    

    def generate_response(self, messages, search_results=None):

        """

        Generate a response based on conversation history.

        Optionally uses search results if provided.

        """

        # Format the conversation

        prompt = self.format_conversation(messages, search_results)

        

        # Generate response

        outputs = self.generator(

            prompt,

            max_new_tokens=150,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        # Extract the generated text

        generated_text = outputs[0]['generated_text']

        

        # Remove the prompt part

        response = generated_text[len(prompt):].strip()

        

        # Clean up

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


We modified the format_conversation and generate_response methods to accept optional search_results. When search results are provided, they are included in the prompt, giving the model additional context to generate better responses.


STEP SIX: BUILDING THE SIDEBAR WITH CONTROLS

Now let us add a sidebar to our interface with controls for search mode and session management. The sidebar will appear on the left side of the screen and contain buttons and toggles.

Add this code to create the sidebar:


# Sidebar for controls

with st.sidebar:

    st.header("⚙️ Settings")

    

    # Search toggle

    search_enabled = st.toggle(

        "🔍 Enable Internet Search",

        value=st.session_state.search_enabled,

        help="When enabled, the chatbot will search the internet for current information"

    )

    st.session_state.search_enabled = search_enabled

    

    # Display current status

    if search_enabled:

        st.success("Search Mode: ON")

        st.info("The chatbot will search the internet before responding")

    else:

        st.info("Search Mode: OFF")

        st.info("The chatbot will use only its training data")

    

    st.divider()

    

    # Session information

    st.header("📝 Session Info")

    st.text(f"Session ID: {st.session_state.session_id}")

    st.text(f"Messages: {len(st.session_state.messages)}")

    

    # Button to start new session

    if st.button("🔄 New Session"):

        save_conversation()

        st.session_state.messages = []

        st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")

        st.rerun()

    

    st.divider()

    

    # Button to save conversation

    if st.button("đź’ľ Save Conversation"):

        save_conversation()

        st.success("Conversation saved!")

    

    # Button to load previous conversations

    st.header("📚 Previous Sessions")

    conversation_files = get_saved_conversations()

    

    if conversation_files:

        selected_file = st.selectbox(

            "Load a previous conversation:",

            conversation_files

        )

        

        if st.button("đź“‚ Load Selected"):

            load_conversation(selected_file)

            st.rerun()

    else:

        st.text("No saved conversations yet")


This sidebar includes a toggle for search mode, displays session information, and provides buttons for managing conversations. The toggle lets users turn internet search on or off. The session info shows the unique ID and message count. Users can start a new session, which saves the current one and clears the screen. They can also manually save conversations or load previous ones.


STEP SEVEN: IMPLEMENTING SESSION PERSISTENCE

To save and load conversations, we need functions that write to and read from files. Let us create a directory for storing conversations and implement the necessary functions:


# Create directory for saved conversations if it doesn't exist

SAVE_DIR = "saved_conversations"

if not os.path.exists(SAVE_DIR):

    os.makedirs(SAVE_DIR)


Add these functions:


def save_conversation():

    """

    Save the current conversation to a JSON file.

    """

    if len(st.session_state.messages) == 0:

        return

    

    filename = f"{st.session_state.session_id}.json"

    filepath = os.path.join(SAVE_DIR, filename)

    

    conversation_data = {

        "session_id": st.session_state.session_id,

        "timestamp": datetime.now().isoformat(),

        "messages": st.session_state.messages

    }

    

    with open(filepath, 'w') as f:

        json.dump(conversation_data, f, indent=2)


def get_saved_conversations():

    """

    Get a list of all saved conversation files.

    """

    if not os.path.exists(SAVE_DIR):

        return []

    

    files = [f for f in os.listdir(SAVE_DIR) if f.endswith('.json')]

    # Sort by modification time, most recent first

    files.sort(key=lambda x: os.path.getmtime(os.path.join(SAVE_DIR, x)), reverse=True)

    return files


def load_conversation(filename):

    """

    Load a conversation from a saved file.

    """

    filepath = os.path.join(SAVE_DIR, filename)

    

    with open(filepath, 'r') as f:

        conversation_data = json.load(f)

    

    st.session_state.messages = conversation_data["messages"]

    st.session_state.session_id = conversation_data["session_id"]


These functions handle saving conversations as JSON files and loading them back. Each conversation is saved with its session ID as the filename. The JSON file contains the session ID, timestamp, and all messages. When loading, we restore the messages and session ID to session state.


STEP EIGHT: PUTTING IT ALL TOGETHER

Now we need to modify our chat input section to use all the components we have built. Replace the simple echo code with this complete implementation:


# Chat input

if prompt := st.chat_input("Type your message here..."):

    # Add user message to chat

    st.session_state.messages.append({"role": "user", "content": prompt})

    with st.chat_message("user"):

        st.markdown(prompt)

    

    # Generate response

    with st.chat_message("assistant"):

        with st.spinner("Thinking..."):

            # Check if search is enabled

            if st.session_state.search_enabled:

                # Perform internet search

                with st.status("Searching the internet...", expanded=True):

                    st.write("🔍 Performing web search...")

                    search_results = search_internet(prompt)

                    st.write("✅ Search complete!")

                

                # Generate response with search results

                response = conversation_manager.generate_response(

                    st.session_state.messages,

                    search_results=search_results

                )

            else:

                # Generate response without search

                response = conversation_manager.generate_response(

                    st.session_state.messages

                )

            

            # Display the response

            st.markdown(response)

    

    # Add assistant response to chat history

    st.session_state.messages.append({"role": "assistant", "content": response})

    

    # Auto-save after each interaction

    save_conversation()


This is where everything comes together. When a user sends a message, we first add it to our message history and display it. Then we check if search mode is enabled. If it is, we perform a web search and show the progress with status messages. We then generate a response using either just the conversation history or the conversation history plus search results. Finally, we display the response, add it to our history, and automatically save the conversation.


STEP NINE: ADDING ERROR HANDLING AND POLISH

Let us add some error handling and polish to make our chatbot more robust. Wrap the main chat functionality in a try-except block:


# Main chat interface

try:

    # Chat input

    if prompt := st.chat_input("Type your message here..."):

        # Add user message

        st.session_state.messages.append({"role": "user", "content": prompt})

        with st.chat_message("user"):

            st.markdown(prompt)

        

        # Generate response

        with st.chat_message("assistant"):

            try:

                with st.spinner("Thinking..."):

                    if st.session_state.search_enabled:

                        with st.status("Searching the internet...", expanded=True):

                            st.write("🔍 Performing web search...")

                            search_results = search_internet(prompt)

                            st.write("✅ Search complete!")

                        

                        response = conversation_manager.generate_response(

                            st.session_state.messages,

                            search_results=search_results

                        )

                    else:

                        response = conversation_manager.generate_response(

                            st.session_state.messages

                        )

                    

                    st.markdown(response)

            

            except Exception as e:

                error_msg = f"Sorry, I encountered an error: {str(e)}"

                st.error(error_msg)

                response = error_msg

        

        # Add response to history

        st.session_state.messages.append({"role": "assistant", "content": response})

        

        # Auto-save

        save_conversation()


except Exception as e:

    st.error(f"An error occurred: {str(e)}")


This error handling ensures that if something goes wrong, the app will not crash. Instead, it will show an error message to the user.

Let us also add a footer with helpful information:


# Footer

st.divider()

st.markdown("""

### đź’ˇ Tips for Using This Chatbot:

- Toggle **Internet Search** ON to get current information from the web

- The chatbot remembers your conversation history during each session

- Click **New Session** to start fresh (your current conversation will be saved)

- Save important conversations using the **Save** button

- Load previous conversations from the sidebar

""")


THE COMPLETE WORKING CODE

Here is the complete code for our chatbot application. Save this as “chatbot_app.py”:


import streamlit as st

import torch

from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline

from langchain_community.tools import DuckDuckGoSearchRun

from datetime import datetime

import json

import os


# Page configuration

st.set_page_config(

    page_title="My AI Chatbot",

    page_icon="🤖",

    layout="wide"

)


# Create directory for saved conversations

SAVE_DIR = "saved_conversations"

if not os.path.exists(SAVE_DIR):

    os.makedirs(SAVE_DIR)


# Device detection function

def get_device():

    """Detect the best available device for running the model."""

    if torch.cuda.is_available():

        device = torch.device("cuda")

        device_name = torch.cuda.get_device_name(0)

        print(f"Using CUDA device: {device_name}")

        return device

    elif torch.backends.mps.is_available():

        device = torch.device("mps")

        print("Using Apple MPS device")

        return device

    else:

        device = torch.device("cpu")

        print("Using CPU device")

        return device


# Get device

device = get_device()


# Load model function

@st.cache_resource

def load_model():

    """Load the language model and tokenizer."""

    model_name = "gpt2"

    

    tokenizer = AutoTokenizer.from_pretrained(model_name)

    model = AutoModelForCausalLM.from_pretrained(

        model_name,

        torch_dtype=torch.float16 if device.type != "cpu" else torch.float32

    )

    model = model.to(device)

    

    generator = pipeline(

        "text-generation",

        model=model,

        tokenizer=tokenizer,

        device=device.type if device.type != "mps" else -1

    )

    

    return generator, tokenizer


# Load search tool

@st.cache_resource

def load_search_tool():

    """Initialize the DuckDuckGo search tool."""

    search = DuckDuckGoSearchRun()

    return search


# Conversation Manager class

class ConversationManager:

    """Manages conversation history and generates responses."""

    

    def __init__(self, generator, tokenizer, max_history=5):

        self.generator = generator

        self.tokenizer = tokenizer

        self.max_history = max_history

    

    def format_conversation(self, messages, search_results=None):

        """Format conversation history into a prompt."""

        recent_messages = messages[-self.max_history:]

        conversation = ""

        

        if search_results:

            conversation += f"Reference Information:\n{search_results}\n\n"

        

        for msg in recent_messages:

            if msg["role"] == "user":

                conversation += f"User: {msg['content']}\n"

            else:

                conversation += f"Assistant: {msg['content']}\n"

        

        conversation += "Assistant:"

        return conversation

    

    def generate_response(self, messages, search_results=None):

        """Generate a response based on conversation history."""

        prompt = self.format_conversation(messages, search_results)

        

        outputs = self.generator(

            prompt,

            max_new_tokens=150,

            num_return_sequences=1,

            temperature=0.7,

            do_sample=True,

            pad_token_id=self.tokenizer.eos_token_id

        )

        

        generated_text = outputs[0]['generated_text']

        response = generated_text[len(prompt):].strip()

        

        if '\n' in response:

            response = response.split('\n')[0]

        

        return response


# Search function

def search_internet(query):

    """Perform an internet search and return results."""

    try:

        results = search_tool.run(query)

        formatted_results = f"Search results for '{query}':\n\n{results}"

        return formatted_results

    except Exception as e:

        return f"Search failed: {str(e)}"


# Conversation saving functions

def save_conversation():

    """Save the current conversation to a JSON file."""

    if len(st.session_state.messages) == 0:

        return

    

    filename = f"{st.session_state.session_id}.json"

    filepath = os.path.join(SAVE_DIR, filename)

    

    conversation_data = {

        "session_id": st.session_state.session_id,

        "timestamp": datetime.now().isoformat(),

        "messages": st.session_state.messages

    }

    

    with open(filepath, 'w') as f:

        json.dump(conversation_data, f, indent=2)


def get_saved_conversations():

    """Get list of all saved conversation files."""

    if not os.path.exists(SAVE_DIR):

        return []

    

    files = [f for f in os.listdir(SAVE_DIR) if f.endswith('.json')]

    files.sort(key=lambda x: os.path.getmtime(os.path.join(SAVE_DIR, x)), reverse=True)

    return files


def load_conversation(filename):

    """Load a conversation from a saved file."""

    filepath = os.path.join(SAVE_DIR, filename)

    

    with open(filepath, 'r') as f:

        conversation_data = json.load(f)

    

    st.session_state.messages = conversation_data["messages"]

    st.session_state.session_id = conversation_data["session_id"]


# Load model and tools

with st.spinner("Loading AI model... This may take a minute..."):

    generator, tokenizer = load_model()

    search_tool = load_search_tool()


# Initialize conversation manager

conversation_manager = ConversationManager(generator, tokenizer)


# Initialize session state

if 'messages' not in st.session_state:

    st.session_state.messages = []


if 'search_enabled' not in st.session_state:

    st.session_state.search_enabled = False


if 'session_id' not in st.session_state:

    st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")


# Title

st.title("🤖 My AI Chatbot")

st.markdown("Chat with an AI powered by Large Language Models")


# Sidebar

with st.sidebar:

    st.header("⚙️ Settings")

    

    search_enabled = st.toggle(

        "🔍 Enable Internet Search",

        value=st.session_state.search_enabled,

        help="When enabled, the chatbot will search the internet for current information"

    )

    st.session_state.search_enabled = search_enabled

    

    if search_enabled:

        st.success("Search Mode: ON")

        st.info("The chatbot will search the internet before responding")

    else:

        st.info("Search Mode: OFF")

        st.info("The chatbot will use only its training data")

    

    st.divider()

    

    st.header("📝 Session Info")

    st.text(f"Session ID: {st.session_state.session_id}")

    st.text(f"Messages: {len(st.session_state.messages)}")

    

    if st.button("🔄 New Session"):

        save_conversation()

        st.session_state.messages = []

        st.session_state.session_id = datetime.now().strftime("%Y%m%d_%H%M%S")

        st.rerun()

    

    st.divider()

    

    if st.button("đź’ľ Save Conversation"):

        save_conversation()

        st.success("Conversation saved!")

    

    st.header("📚 Previous Sessions")

    conversation_files = get_saved_conversations()

    

    if conversation_files:

        selected_file = st.selectbox(

            "Load a previous conversation:",

            conversation_files

        )

        

        if st.button("đź“‚ Load Selected"):

            load_conversation(selected_file)

            st.rerun()

    else:

        st.text("No saved conversations yet")


# Display chat messages

for message in st.session_state.messages:

    with st.chat_message(message["role"]):

        st.markdown(message["content"])


# Chat input

try:

    if prompt := st.chat_input("Type your message here..."):

        st.session_state.messages.append({"role": "user", "content": prompt})

        with st.chat_message("user"):

            st.markdown(prompt)

        

        with st.chat_message("assistant"):

            try:

                with st.spinner("Thinking..."):

                    if st.session_state.search_enabled:

                        with st.status("Searching the internet...", expanded=True):

                            st.write("🔍 Performing web search...")

                            search_results = search_internet(prompt)

                            st.write("✅ Search complete!")

                        

                        response = conversation_manager.generate_response(

                            st.session_state.messages,

                            search_results=search_results

                        )

                    else:

                        response = conversation_manager.generate_response(

                            st.session_state.messages

                        )

                    

                    st.markdown(response)

            

            except Exception as e:

                error_msg = f"Sorry, I encountered an error: {str(e)}"

                st.error(error_msg)

                response = error_msg

        

        st.session_state.messages.append({"role": "assistant", "content": response})

        save_conversation()


except Exception as e:

    st.error(f"An error occurred: {str(e)}")


# Footer

st.divider()

st.markdown("""

### đź’ˇ Tips for Using This Chatbot:

- Toggle **Internet Search** ON to get current information from the web

- The chatbot remembers your conversation history during each session

- Click **New Session** to start fresh (your current conversation will be saved)

- Save important conversations using the **Save** button

- Load previous conversations from the sidebar

""")


RUNNING YOUR CHATBOT

Now that you have created your complete chatbot application, it is time to run it and see it in action. Make sure your virtual environment is still activated. In your terminal, navigate to the directory containing your “chatbot_app.py” file. Then run this command:


streamlit run chatbot_app.py


After a few seconds, your default web browser should automatically open and display your chatbot interface. If it does not open automatically, look at your terminal output. Streamlit will print a local URL like “http://localhost:8501”. You can copy and paste this URL into your browser.

The first time you run the chatbot, it will take a minute or two to load the AI model. You will see a loading message while this happens. Once the model is loaded, you are ready to start chatting.

Try having a conversation with your chatbot. Type a message in the input box at the bottom of the screen and press Enter. The chatbot will think for a moment and then respond. Notice how it remembers previous messages in the conversation - you can refer back to things you discussed earlier and it will understand the context.

Now try enabling internet search by clicking the toggle in the sidebar. Ask a question about something current or recent, like “What are the latest developments in space exploration?” The chatbot will perform a web search, and you will see status messages showing this process. The response should incorporate information from the search results.

Experiment with starting new sessions and loading previous conversations. Each session gets saved automatically, so you never lose your chat history. You can come back days later and resume old conversations.


UNDERSTANDING HOW THE COMPONENTS WORK TOGETHER

Now that you have a working chatbot, let us review how all the pieces fit together to create the complete system. Understanding this architecture will help you customize and extend your chatbot.

The application starts by loading the AI model and search tool. This happens once when the app first starts, thanks to the Streamlit caching decorators. The model is moved to the best available device - CUDA for NVIDIA GPUs, MPS for Apple Silicon, or CPU as a fallback.

When you type a message, Streamlit captures it and adds it to the session state messages list. This list persists across interactions because of how Streamlit’s session state works. Your message is immediately displayed in the chat interface.

Next, the application checks whether internet search is enabled. This is controlled by the toggle button in the sidebar, which updates a session state variable. If search is on, the DuckDuckGo search tool is called with your message as the query. The search results are retrieved and formatted into a string.

The conversation manager then takes over. It looks at the message history and optionally the search results, and formats them into a prompt string that the AI model can understand. This prompt includes recent messages to give context, and the search results if available to provide additional information.

The formatted prompt is sent to the AI model through the text generation pipeline. The model processes this input and generates a response. The generation parameters control aspects like response length and creativity. The model’s raw output includes both the input prompt and the generated response, so we extract just the new text.

This generated response is displayed in the chat interface and added to the message history. Finally, the entire conversation is automatically saved to a JSON file in the saved_conversations directory. This happens after every interaction, so you never lose data even if you close the browser.

The sidebar provides controls that modify this flow. The search toggle changes whether internet lookup happens. The new session button saves the current conversation and resets the message history. The load conversation feature retrieves a previously saved conversation and restores it to the current session.


TESTING YOUR CHATBOT THOROUGHLY

Before you share your chatbot or rely on it for important tasks, you should test it thoroughly to understand its capabilities and limitations. Here are some testing scenarios you should try.

First, test basic conversational abilities. Have a multi-turn conversation where you ask follow-up questions that refer to previous messages. For example, ask “What is Python?”, then “What are its main uses?”, then “Which one would you recommend for a beginner?” The chatbot should maintain context throughout.

Second, test the search functionality. Turn on internet search and ask questions about recent events or current information that would not be in the model’s training data. Compare the responses with search enabled versus disabled to see the difference. Good test questions include “What is the current date?”, “What are today’s top news headlines?”, or “What is the weather in [your city]?”

Third, test the session management. Create a conversation, start a new session, then load the previous conversation to verify it was saved correctly. Try loading multiple different conversations to ensure the system correctly switches between them.

Fourth, test error handling. Try to break the system by entering very long messages, special characters, or unusual inputs. The error handling should catch problems and display user-friendly messages rather than crashing.

Fifth, test hardware acceleration. Look at the console output when the app starts to verify it detected your hardware correctly. If you have a GPU, the model should run faster than on CPU.

Make notes about what works well and what could be improved. Every limitation you discover is an opportunity for enhancement.


TROUBLESHOOTING COMMON ISSUES


As you work with your chatbot, you might encounter some common issues. Here are solutions to the most frequent problems.

If the model takes a very long time to load, this is normal the first time. The Hugging Face Transformers library downloads the model files from the internet, which can take several minutes depending on your connection speed. Subsequent runs will be much faster because the model is cached locally.

If you get an error about not having enough memory, the AI model might be too large for your system. Try using a smaller model by changing the model_name variable to “distilgpt2”, which is a compressed version of GPT-2 that uses less memory. Alternatively, make sure you close other memory-intensive applications.

If the search functionality does not work, verify that you installed the duckduckgo-search package correctly. You can test it separately by opening a Python console and running “from langchain_community.tools import DuckDuckGoSearchRun” followed by “search = DuckDuckGoSearchRun()” and “search.run(‘test query’)”. If this works, the problem is elsewhere in your code.

If saved conversations do not load, check that the saved_conversations directory exists and contains JSON files. You can open these files in a text editor to verify they are formatted correctly. The load_conversation function expects specific keys in the JSON structure.

If responses seem inconsistent or low quality, try adjusting the generation parameters. Lowering the temperature makes responses more focused and predictable. Increasing max_new_tokens allows longer responses. Experiment with different values to find what works best.

If the application crashes with errors about device placement, check that your PyTorch installation matches your hardware. For NVIDIA GPUs, ensure you installed the CUDA-enabled version. For Apple Silicon, verify you have the latest PyTorch version with MPS support.


FUTURE ENHANCEMENTS AND EXTENSIONS

Now that you have a working chatbot, there are many ways you can enhance and extend it. Here are some ideas for improvements you might want to implement.

One significant enhancement would be using a more powerful language model. GPT-2 is good for learning, but larger models produce better responses. You could upgrade to GPT-2 Medium or Large, or try other models like Microsoft’s DialoGPT which is specifically trained for conversations. More advanced users might experiment with models like FLAN-T5 or LLaMA.

You could add support for multiple search engines. Currently we only use DuckDuckGo, but LangChain supports other search tools like Google Search, Wikipedia, and Wolfram Alpha. You could let users choose which search engine to use, or have the chatbot automatically select the best one based on the question type.

Another useful feature would be document upload. You could let users upload PDF or text files, and the chatbot could answer questions about the content. This would require adding document loading and text splitting capabilities, which LangChain provides through its document loaders and text splitters.

You could implement a more sophisticated conversation history system. Instead of just keeping recent messages, you could use embeddings to find the most relevant past messages, or implement summarization to condense long conversations. This would let the chatbot maintain context over much longer conversations without exceeding token limits.

Adding user authentication would let multiple people use the same chatbot instance while keeping their conversations separate and private. You could use Streamlit’s authentication components or integrate with existing authentication systems.

You could add voice input and output. The speech_recognition library can convert speech to text, and text-to-speech libraries like pyttsx3 or gTTS can convert responses to audio. This would create a more natural, hands-free interaction model.

Another enhancement would be adding the ability to generate images using models like Stable Diffusion when users ask for visual content. You could integrate this as another tool the chatbot can use when appropriate.

You could also implement a rating system where users can rate responses, and use this feedback to fine-tune the model or select better generation parameters. This creates a feedback loop for continuous improvement.

For more advanced users, you could implement Retrieval Augmented Generation or RAG, which combines the language model with a vector database of documents. This allows the chatbot to draw from a large knowledge base while still generating natural language responses.

You might add support for different languages by using multilingual models or translation APIs. This would let users chat in their preferred language.

Finally, you could deploy your chatbot to the cloud so others can access it. Streamlit offers Streamlit Cloud for easy deployment, or you could use platforms like Heroku, AWS, or Google Cloud Platform.


BEST PRACTICES FOR LLM APPLICATION DEVELOPMENT

As you continue developing your chatbot and other AI applications, keep these best practices in mind to create high-quality, maintainable systems.

Always handle errors gracefully. AI models can be unpredictable, and external services like search APIs can fail. Wrap potentially problematic code in try-except blocks and provide helpful error messages to users. Never let exceptions crash your application.

Use caching strategically. Loading models and initializing components is expensive. The Streamlit cache_resource decorator ensures these operations happen only once, dramatically improving performance. However, be careful not to cache things that should change between interactions.

Monitor resource usage. AI models consume significant memory and processing power. Limit conversation history length to avoid exceeding memory limits. Consider implementing cleanup mechanisms for long-running applications.

Document your code thoroughly. Even though this tutorial includes extensive comments, in real projects you should add even more documentation. Explain why you made certain design decisions, not just what the code does. Future you will thank present you.

Test with diverse inputs. Users will interact with your chatbot in unexpected ways. Test edge cases, unusual inputs, and potential misuse scenarios. Ensure your application behaves appropriately in all situations.

Be transparent about limitations. AI models can produce incorrect or biased responses. Make sure users understand they are interacting with an AI and should verify important information. Consider adding disclaimers or warnings for sensitive use cases.

Respect privacy and security. If your chatbot handles personal information, ensure you follow data protection best practices. Do not log sensitive information, use encryption for data at rest and in transit, and provide clear privacy policies.

Optimize for user experience. Response time matters. If generation takes too long, users will lose interest. Show progress indicators during long operations. Make the interface intuitive and provide helpful guidance.

Keep models and libraries updated. The AI field evolves rapidly. Regularly update your dependencies to get bug fixes, performance improvements, and new features. However, test thoroughly after updates to ensure nothing breaks.

Learn from user interactions. If possible, collect anonymized usage data to understand how people use your chatbot. This helps identify which features are valuable and where users struggle.


UNDERSTANDING THE ETHICAL IMPLICATIONS

As you develop AI applications, it is important to understand the ethical implications and responsibilities that come with this technology. AI systems can have significant impacts on people’s lives, and developers have a responsibility to create systems that are beneficial and minimize harm.

Large Language Models can sometimes generate biased, incorrect, or harmful content. This happens because they are trained on internet text, which contains human biases and errors. As a developer, you should be aware of these limitations and implement safeguards. Consider adding content filtering to block inappropriate outputs, and be transparent with users about potential issues.

Privacy is a critical concern. Our chatbot saves conversation history, which might contain sensitive information. Make sure users understand what data is collected and how it is used. Implement appropriate security measures to protect this data. Give users control over their data, including the ability to delete their information.

The environmental impact of AI is another consideration. Training large models requires enormous computational resources and energy. While you are not training models in this tutorial, be mindful of energy efficiency in your applications. Using smaller models when possible and optimizing code can reduce environmental impact.

Accessibility is important too. Not everyone interacts with technology in the same way. Consider how users with disabilities might use your chatbot. Can it be used with screen readers? Can users adjust text size? Thinking about accessibility from the start makes your application more inclusive.

Be honest about what your chatbot can and cannot do. Do not overstate its capabilities or present it as something it is not. If it is a simple model with limited abilities, be clear about that. Users should understand they are interacting with an AI, not a human expert.

Think about potential misuse. Could your chatbot be used to spread misinformation, create spam, or harm others? While you cannot prevent all misuse, you can design your system to make harmful uses more difficult. This might include rate limiting, content filtering, or requiring authentication.


LEARNING RESOURCES FOR CONTINUED GROWTH

This tutorial has given you a solid foundation in building LLM applications, but there is much more to learn. Here are resources to help you continue your learning journey.

The official documentation for the libraries we used is essential reading. The Streamlit documentation at docs.streamlit.io provides detailed guides and examples for building web applications. The LangChain documentation at python.langchain.com covers all aspects of the framework, including advanced features we did not explore. The Hugging Face Transformers documentation at huggingface.co/docs/transformers is comprehensive and includes tutorials for many different tasks.

Online courses can provide structured learning paths. Coursera and Udemy offer courses on natural language processing, deep learning, and LLM applications. Many are taught by university professors and industry experts. Some are free, others require payment but often provide certificates.

YouTube has excellent tutorial channels. Channels like “Corey Schafer” and “Tech With Tim” have Python tutorials. For AI-specific content, “AssemblyAI” and “1littlecoder” create excellent tutorials on building AI applications. “Hugging Face” has their own channel with model demonstrations and technical talks.

Reading research papers, while challenging, helps you understand the latest developments. Papers like “Attention Is All You Need” which introduced Transformers, and GPT papers from OpenAI provide insights into how these models work. The ArXiv website hosts pre-publication papers on AI and machine learning.

Books provide in-depth knowledge. “Natural Language Processing with Transformers” by Lewis Tunstall and others covers modern NLP techniques. “Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” by AurĂ©lien GĂ©ron is an excellent introduction to machine learning. “Deep Learning” by Ian Goodfellow is more advanced but comprehensive.

Online communities are valuable for getting help and sharing knowledge. The Streamlit forum, LangChain Discord server, and Hugging Face forums are active communities where you can ask questions and learn from others. Reddit communities like r/MachineLearning and r/LanguageTechnology discuss AI developments and applications.

GitHub is an incredible resource for learning from real projects. Search for repositories using the technologies we covered. Reading other people’s code is one of the best ways to learn new techniques and best practices. You can also contribute to open source projects to gain experience.

Experiment with different models and techniques. The best way to learn is by doing. Try modifying your chatbot to add new features. Experiment with different models from Hugging Face. Build completely new applications using what you have learned. Each project will teach you something new.


CONCLUSION

Congratulations on building your own AI chatbot. You have learned a tremendous amount in this tutorial, from understanding Large Language Models to implementing a complete application with a web interface, conversation management, internet search capabilities, and hardware acceleration.

The chatbot you built is not just a learning exercise - it is a functional application that demonstrates real AI capabilities. You have conversation history that maintains context across multiple exchanges. You have session management that saves and loads conversations. You have internet search that provides up-to-date information. And you have hardware detection that ensures your application runs efficiently on different systems.

More importantly, you have learned the fundamental concepts and techniques used in modern AI application development. You understand how to work with pre-trained models, how to build user interfaces with Streamlit, how to use frameworks like LangChain to simplify complex tasks, and how to integrate multiple components into a cohesive system.

The skills you have learned are applicable far beyond this specific chatbot. The patterns and techniques we used - component-based architecture, session management, error handling, and hardware optimization - apply to many types of applications. The experience of reading documentation, debugging problems, and iterating on features is essential to all software development.

As you continue your journey in AI and software development, remember that learning is a continuous process. Technology evolves rapidly, especially in AI. New models, frameworks, and techniques emerge constantly. Stay curious, keep experimenting, and never stop learning.

Your chatbot is just the beginning. With the foundation you have built, you can create increasingly sophisticated applications. Maybe you will build a chatbot that helps students learn difficult subjects. Or a virtual assistant that helps people manage their daily tasks. Or a research tool that helps scientists explore complex questions. The possibilities are limited only by your imagination and dedication.

Thank you for working through this tutorial. We hope you found it educational, engaging, and inspiring. Now go forth and build amazing things with AI.


(to be CONTINUED)


ADDENDUM: RESOURCES


BOOKS FOR FURTHER LEARNING

“Natural Language Processing with Transformers” by Lewis Tunstall, Leandro von Werra, and Thomas Wolf. Published by O’Reilly Media. This book is specifically focused on using Transformers for NLP tasks and covers practical applications. It is very hands-on and includes numerous code examples using the Hugging Face library.

“Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” by AurĂ©lien GĂ©ron. Published by O’Reilly Media. While not specifically about LLMs, this book provides essential machine learning fundamentals that underpin AI applications. The second edition covers deep learning extensively.

“Deep Learning” by Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Published by MIT Press. This is a comprehensive textbook that covers the mathematical and theoretical foundations of deep learning. It is more academic than the others but provides deep understanding.

“Building Machine Learning Powered Applications” by Emmanuel Ameisen. Published by O’Reilly Media. This book focuses on the practical aspects of deploying machine learning models in production, including many considerations we touched on in this tutorial.

“Speech and Language Processing” by Dan Jurafsky and James H. Martin. Available free online. This is a comprehensive textbook on natural language processing that covers both traditional and modern approaches. The third edition includes coverage of neural approaches.


WEBSITES AND ONLINE DOCUMENTATION

Official Streamlit Documentation at https://docs.streamlit.io provides comprehensive guides, API references, and examples for building web applications with Streamlit.

LangChain Documentation at https://python.langchain.com/docs/get_started/introduction covers the entire LangChain framework with tutorials, how-to guides, and conceptual explanations.

Hugging Face Documentation at https://huggingface.co/docs provides documentation for all Hugging Face libraries including Transformers, Datasets, and Accelerate. The Hugging Face Hub at https://huggingface.co/models hosts thousands of pre-trained models.

PyTorch Documentation at https://pytorch.org/docs/stable/index.html covers the PyTorch deep learning framework in detail, including guides for using CUDA and MPS acceleration.

The official Python documentation at https://docs.python.org/3/ is an essential reference for Python itself.


ONLINE COURSES AND TUTORIALS

Coursera offers several relevant courses. “Natural Language Processing” by DeepLearning.AI covers modern NLP techniques. “Deep Learning Specialization” by Andrew Ng is a comprehensive introduction to deep learning. “Machine Learning” by Andrew Ng is a classic introduction to machine learning fundamentals.

Fast.ai offers free courses including “Practical Deep Learning for Coders” which takes a code-first approach to learning deep learning. The course includes a chapter on natural language processing.

The Hugging Face Course at https://huggingface.co/course/chapter1/1 is a free comprehensive course on using Transformers for various NLP tasks. It includes hands-on exercises and covers advanced topics.

YouTube has many excellent channels. “3Blue1Brown” explains mathematical concepts visually, which is helpful for understanding how neural networks work. “Sentdex” has practical Python and machine learning tutorials. “Two Minute Papers” summarizes recent AI research papers in accessible formats.


RESEARCH PAPERS AND TECHNICAL BLOGS

“Attention Is All You Need” by Vaswani et al. introduced the Transformer architecture that powers modern LLMs. Available on ArXiv at https://arxiv.org/abs/1706.03762

“Language Models are Few-Shot Learners” describes GPT-3 and demonstrates the capabilities of large language models. Available at https://arxiv.org/abs/2005.14165

The OpenAI Blog at https://openai.com/blog publishes accessible explanations of their research and models.

The Hugging Face Blog at https://huggingface.co/blog features technical tutorials and announcements about new models and techniques.

Google AI Blog at https://ai.googleblog.com discusses Google’s AI research including work on language models.

Anthropic’s Research page at https://www.anthropic.com/research publishes papers about AI safety and capabilities.


COMMUNITIES AND FORUMS

The Streamlit Forum at https://discuss.streamlit.io is an active community where you can ask questions about Streamlit and share your projects.

The LangChain GitHub Discussions at https://github.com/langchain-ai/langchain/discussions is where the LangChain community discusses the framework.

Hugging Face Forums at https://discuss.huggingface.co covers topics related to models, datasets, and the Transformers library.

Reddit has several relevant communities. r/MachineLearning discusses machine learning research and applications. r/LanguageTechnology focuses on NLP and computational linguistics. r/learnmachinelearning is for beginners asking questions and sharing resources.

Discord servers exist for many AI frameworks and tools. The LangChain Discord and Hugging Face Discord are particularly active.

Stack Overflow at https://stackoverflow.com is invaluable for getting help with specific programming problems.


TOOLS AND PLATFORMS

Google Colab at https://colab.research.google.com provides free access to GPUs for running machine learning code in Jupyter notebooks.

Kaggle at https://www.kaggle.com offers free GPU access, datasets, and competitions. It is a good place to practice machine learning.

GitHub at https://github.com is essential for version control and for exploring open source projects. Learning Git is valuable for any programmer.

Weights and Biases at https://wandb.ai provides tools for tracking machine learning experiments.

Streamlit Cloud at https://streamlit.io/cloud offers free deployment for Streamlit applications.