Saturday, October 03, 2026

LANGGRAPH TUTORIAL FOR DEVELOPERS Building Stateful Multi-Agent LLM Applications from Scratch



INTRODUCTION: UNDERSTANDING LANGGRAPH AND ITS PURPOSE

When you begin working with Large Language Models, you quickly discover that single LLM calls are insufficient for complex tasks. Real applications require multiple coordinated steps, decision-making capabilities, state management across interactions, and often multiple specialized agents working together toward a common goal.

LangGraph addresses these challenges by providing a framework for building stateful, multi-agent applications with LLMs. The core insight is that complex LLM workflows can be elegantly modeled as directed graphs, where nodes represent operations or agents, and edges define the flow of information and control between them.

Consider a research assistant application. Such an assistant needs to search for information, analyze findings, synthesize results, and potentially iterate based on what it discovers. Each of these steps might involve different LLM calls with different prompts, external tool usage, and decision points about what to do next. LangGraph provides the structure to orchestrate all these components in a clear, maintainable way.

The graph-based approach offers several advantages. First, it makes your application's logic explicit and visual. You can literally draw out how your application works. Second, it provides fine-grained control over execution flow, allowing you to implement sophisticated conditional logic. Third, it manages state automatically, ensuring that information flows correctly between different parts of your application.

CORE CONCEPTS: THE FOUNDATION OF LANGGRAPH

Before writing any code, we need to understand four fundamental concepts that form the foundation of every LangGraph application. These concepts work together to create powerful, flexible LLM workflows.

The first concept is State. In LangGraph, state represents the shared information that flows through your entire application. Think of it as a living document that every part of your application can read from and write to. As your application executes, moving from one node to another, the state accumulates information, building up context and maintaining the history of what has happened.

State is crucial because it allows different parts of your application to build upon each other's work. When one agent performs a web search, it stores the results in the state. When another agent needs to analyze those results, it can access them from the state. This shared memory is what enables sophisticated multi-step reasoning.

The second concept is the Graph itself. A graph in LangGraph is the overall structure of your application. It defines all the possible paths your application can take and all the operations it can perform. The graph is composed of nodes connected by edges, forming a directed flow of execution.

The third concept is Nodes. Each node in your graph represents a discrete unit of work. A node is implemented as a Python function that receives the current state, performs some operation, and returns updates to the state. The operation could be calling an LLM, invoking an external API, performing calculations, making decisions, or any other computational task your application requires.

The fourth concept is Edges. Edges define how execution flows from one node to another. LangGraph supports two types of edges. Normal edges create unconditional connections, meaning after node A completes, node B always executes next. Conditional edges enable decision-making, allowing your application to choose different paths based on the current state.

ENVIRONMENT SETUP: PREPARING YOUR DEVELOPMENT ENVIRONMENT

To begin working with LangGraph, you need to set up your Python environment with the necessary dependencies. LangGraph requires Python version 3.9 or higher. You will also need to install LangChain, as LangGraph builds upon its foundation.

Open your terminal and execute the following command to install the required packages:

pip install langgraph langchain langchain-openai

For this tutorial, we will use OpenAI's models, so you need an OpenAI API key. You can obtain one from the OpenAI platform website. Once you have your key, set it as an environment variable. On Linux or macOS, use this command:

export OPENAI_API_KEY='your-actual-api-key-here'

On Windows, use this command instead:

set OPENAI_API_KEY=your-actual-api-key-here

Now let's verify that everything is installed correctly with a simple test:

from langgraph.graph import StateGraph, END
from typing import TypedDict, Annotated
import operator

print("LangGraph is successfully installed and ready to use!")

This code imports the essential components we will use throughout this tutorial. The StateGraph class is the primary tool for constructing graphs. The END constant is a special marker indicating workflow completion. The TypedDict and Annotated types from Python's typing module help us define type-safe state structures.

DEFINING STATE: THE INFORMATION BACKBONE

State definition is the first step in building any LangGraph application. The state structure determines what information your application tracks and how that information is updated as the application executes.

LangGraph uses Python's TypedDict to define state schemas. This provides type safety and makes your code more maintainable. Let's start with a simple example:

from typing import TypedDict, Annotated
import operator

class SimpleState(TypedDict):
    counter: int
    message: str

This SimpleState definition creates a state structure with two fields. The counter field stores an integer value, and the message field stores a string. When a node updates these fields, it simply replaces the old value with the new value.

However, LangGraph offers a more sophisticated mechanism for state updates through the Annotated type. This allows you to specify how updates should be applied. The most common pattern uses operator.add to accumulate values rather than replace them:

from typing import TypedDict, Annotated, Sequence
import operator

class ConversationState(TypedDict):
    messages: Annotated[Sequence[str], operator.add]
    step_count: int
    current_topic: str

In this ConversationState definition, the messages field uses Annotated with operator.add. This tells LangGraph that when a node returns new messages, they should be appended to the existing messages list rather than replacing it. This is essential for maintaining conversation history.

The step_count and current_topic fields do not use Annotated, so they follow the default behavior of replacement. When a node returns a new value for step_count, it overwrites the previous value.

Let's see a more complex state definition that might be used for a research assistant:

from typing import TypedDict, Annotated, Sequence, Optional
import operator

class ResearchState(TypedDict):
    # Accumulate messages throughout the conversation
    messages: Annotated[Sequence[str], operator.add]
    # Store search queries that have been executed
    search_queries: Annotated[list[str], operator.add]
    # Store search results from external sources
    search_results: Annotated[list[dict], operator.add]
    # Current research question being investigated
    current_question: str
    # Final synthesized answer
    final_answer: Optional[str]
    # Number of research iterations performed
    iteration_count: int

This ResearchState demonstrates a realistic state structure for a multi-step research application. The messages, search_queries, and search_results fields all use operator.add to accumulate information over time. The current_question, final_answer, and iteration_count fields use replacement semantics.

Understanding how state updates work is critical. When a node function returns a dictionary, LangGraph merges that dictionary into the current state. For fields annotated with operator.add, the new values are added to existing values. For other fields, the new values replace the old values.

CREATING NODES: THE WORKHORSES OF YOUR APPLICATION

Nodes are where the actual work happens in your LangGraph application. Each node is a Python function that takes the current state as input and returns a dictionary representing updates to that state.

Let's create a simple node that increments a counter:

def increment_counter_node(state: SimpleState) -> dict:
    """
    This node increments the counter in the state by one.
    It demonstrates the basic pattern of reading from state
    and returning an update.
    """
    current_count = state.get("counter", 0)
    new_count = current_count + 1
    
    print(f"Incrementing counter from {current_count} to {new_count}")
    
    # Return a dictionary with the updates to apply to state
    return {"counter": new_count}

This increment_counter_node function demonstrates the fundamental node pattern. It receives the state, extracts the current counter value, increments it, and returns a dictionary containing the new counter value. LangGraph automatically merges this update into the state.

Now let's create a more sophisticated node that calls an LLM:

from langchain_openai import ChatOpenAI
from langchain.schema import HumanMessage, AIMessage

def llm_response_node(state: ConversationState) -> dict:
    """
    This node calls an LLM with the current conversation history
    and returns the LLM's response, which gets added to the messages.
    """
    # Initialize the language model
    llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
    
    # Get the current messages from state
    current_messages = state.get("messages", [])
    
    # Convert string messages to LangChain message objects
    formatted_messages = []
    for i, msg in enumerate(current_messages):
        if i % 2 == 0:
            formatted_messages.append(HumanMessage(content=msg))
        else:
            formatted_messages.append(AIMessage(content=msg))
    
    # Call the LLM
    response = llm.invoke(formatted_messages)
    
    print(f"LLM responded: {response.content[:100]}...")
    
    # Return the new message to be added to the conversation
    # Because messages uses operator.add, this will be appended
    return {
        "messages": [response.content],
        "step_count": state.get("step_count", 0) + 1
    }

This llm_response_node demonstrates a more complex operation. It retrieves the conversation history from state, formats it appropriately for the LLM, invokes the LLM, and returns both the new message and an updated step count. Notice how the function returns a dictionary with updates for multiple state fields.

Let's create another node that performs a simulated web search:

def web_search_node(state: ResearchState) -> dict:
    """
    This node simulates performing a web search based on
    the current research question and stores the results.
    """
    question = state.get("current_question", "")
    
    print(f"Performing web search for: {question}")
    
    # In a real application, this would call an actual search API
    # For demonstration, we'll create simulated results
    simulated_results = [
        {
            "title": f"Result 1 for {question}",
            "snippet": "This is a simulated search result snippet...",
            "url": "https://example.com/result1"
        },
        {
            "title": f"Result 2 for {question}",
            "snippet": "Another simulated search result snippet...",
            "url": "https://example.com/result2"
        }
    ]
    
    # Return updates to state
    # search_queries and search_results use operator.add, so these append
    return {
        "search_queries": [question],
        "search_results": simulated_results
    }

This web_search_node shows how you might integrate external tools into your LangGraph application. The node reads the current question from state, performs an operation (in this case simulated, but in reality would call an actual search API), and returns the results to be accumulated in the state.

BUILDING YOUR FIRST GRAPH: PUTTING IT ALL TOGETHER

Now that we understand state and nodes, let's build our first complete LangGraph application. We'll create a simple conversation system that takes user input, processes it through an LLM, and returns a response.

First, let's define our state:

from typing import TypedDict, Annotated, Sequence
import operator

class ChatState(TypedDict):
    messages: Annotated[Sequence[str], operator.add]
    conversation_active: bool

Next, let's create the nodes we'll need:

from langchain_openai import ChatOpenAI
from langchain.schema import HumanMessage, SystemMessage

def chat_node(state: ChatState) -> dict:
    """
    This node processes the conversation through an LLM.
    """
    llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
    
    messages = state.get("messages", [])
    
    # Create a system message to set context
    system_msg = SystemMessage(
        content="You are a helpful assistant. Provide clear, concise answers."
    )
    
    # Format the conversation history
    formatted_messages = [system_msg]
    for i, msg in enumerate(messages):
        formatted_messages.append(HumanMessage(content=msg))
    
    # Get LLM response
    response = llm.invoke(formatted_messages)
    
    print(f"Assistant: {response.content}")
    
    return {
        "messages": [response.content]
    }

Let's build the graph itself:

from langgraph.graph import StateGraph, END

def create_simple_chat_graph():
    """
    Creates a simple chat graph with one LLM node.
    """
    # Initialize the graph with our state type
    workflow = StateGraph(ChatState)
    
    # Add the chat node to the graph
    # First argument is the node name, second is the function
    workflow.add_node("chat", chat_node)
    
    # Set the entry point - where execution begins
    workflow.set_entry_point("chat")
    
    # Add an edge from chat node to END
    # This means after the chat node executes, the workflow terminates
    workflow.add_edge("chat", END)
    
    # Compile the graph into an executable application
    app = workflow.compile()
    
    return app

This create_simple_chat_graph function demonstrates the basic pattern for building graphs. We create a StateGraph instance, add our nodes, define the entry point, connect nodes with edges, and compile the graph into an executable application.

Let's use this graph:

# Create the graph application
app = create_simple_chat_graph()

# Prepare initial state with a user message
initial_state = {
    "messages": ["What is LangGraph and why is it useful?"],
    "conversation_active": True
}

# Execute the graph
result = app.invoke(initial_state)

# The result contains the final state after execution
print("\nFinal conversation:")
for i, msg in enumerate(result["messages"]):
    role = "User" if i % 2 == 0 else "Assistant"
    print(f"{role}: {msg}")

When you run this code, the graph executes the chat node, which processes the user's question through the LLM and returns a response. The final state contains both the original user message and the assistant's response.

CONDITIONAL EDGES: MAKING INTELLIGENT DECISIONS

The real power of LangGraph emerges when you add conditional logic to your graphs. Conditional edges allow your application to make decisions about which node to execute next based on the current state.

To implement conditional edges, you create a router function that examines the state and returns the name of the next node to execute. Let's build an example that demonstrates this:

from typing import TypedDict, Annotated, Sequence, Literal
import operator

class TaskState(TypedDict):
    messages: Annotated[Sequence[str], operator.add]
    task_type: str
    task_complete: bool
    result: str

Next let's create a router function:

def route_based_on_task(state: TaskState) -> Literal["math_task", "text_task", "end"]:
    """
    This router function examines the state and decides which node
    should execute next based on the task type and completion status.
    """
    # If task is complete, end the workflow
    if state.get("task_complete", False):
        return "end"
    
    # Otherwise, route based on task type
    task_type = state.get("task_type", "")
    
    if "math" in task_type.lower() or "calculate" in task_type.lower():
        return "math_task"
    else:
        return "text_task"

This router function demonstrates decision-making logic. It checks if the task is complete, and if so, returns "end" to terminate the workflow. Otherwise, it examines the task type and routes to either a math-specialized node or a text-specialized node.

Let's create the specialized nodes:

def math_task_node(state: TaskState) -> dict:
    """
    This node handles mathematical tasks.
    """
    print("Processing mathematical task...")
    
    messages = state.get("messages", [])
    last_message = messages[-1] if messages else ""
    
    # In a real application, this would use an LLM or calculation engine
    result = f"Mathematical analysis of: {last_message}"
    
    return {
        "messages": [result],
        "task_complete": True,
        "result": result
    }

def text_task_node(state: TaskState) -> dict:
    """
    This node handles text-based tasks.
    """
    print("Processing text task...")
    
    messages = state.get("messages", [])
    last_message = messages[-1] if messages else ""
    
    # In a real application, this would use an LLM
    result = f"Text analysis of: {last_message}"
    
    return {
        "messages": [result],
        "task_complete": True,
        "result": result
    }

Now let's build a graph that uses conditional routing:

from langgraph.graph import StateGraph, END

def create_conditional_graph():
    """
    Creates a graph with conditional routing based on task type.
    """
    workflow = StateGraph(TaskState)
    
    # Add both specialized nodes
    workflow.add_node("math_task", math_task_node)
    workflow.add_node("text_task", text_task_node)
    
    # Set entry point to a router
    # We need to add a node that determines the initial route
    workflow.set_entry_point("math_task")
    
    # Add conditional edges from each task node
    # The router function determines where to go next
    workflow.add_conditional_edges(
        "math_task",
        route_based_on_task,
        {
            "end": END,
            "math_task": "math_task",
            "text_task": "text_task"
        }
    )
    
    workflow.add_conditional_edges(
        "text_task",
        route_based_on_task,
        {
            "end": END,
            "math_task": "math_task",
            "text_task": "text_task"
        }
    )
    
    app = workflow.compile()
    return app

The add_conditional_edges method is crucial here. It takes three arguments. First, the source node name. Second, the router function that makes the decision. Third, a mapping dictionary that maps the router's return values to actual node names or END.

Let's test this conditional graph:

app = create_conditional_graph()

# Test with a math task
math_state = {
    "messages": ["Calculate the sum of 15 and 27"],
    "task_type": "math calculation",
    "task_complete": False,
    "result": ""
}

result = app.invoke(math_state)
print(f"Math task result: {result['result']}")

# Test with a text task
text_state = {
    "messages": ["Summarize the benefits of exercise"],
    "task_type": "text summary",
    "task_complete": False,
    "result": ""
}

result = app.invoke(text_state)
print(f"Text task result: {result['result']}")

This example demonstrates how conditional edges enable your application to dynamically choose different execution paths based on the current state, making your LLM applications much more flexible and intelligent.

WORKING WITH LANGCHAIN MESSAGES: PROPER MESSAGE HANDLING

In real LangGraph applications, you'll typically work with LangChain's message types rather than plain strings. LangChain provides several message classes that represent different roles in a conversation.

Let's update our state definition to use proper message types:

from typing import TypedDict, Annotated, Sequence
from langchain.schema import BaseMessage, HumanMessage, AIMessage, SystemMessage
import operator

class ProperChatState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    iteration_count: int

The BaseMessage type is the parent class for all message types in LangChain. Using this in our state definition allows us to store any type of message (HumanMessage, AIMessage, SystemMessage, etc.) in our messages list.

Next let's create a node that properly handles these message types:

from langchain_openai import ChatOpenAI

def proper_chat_node(state: ProperChatState) -> dict:
    """
    This node demonstrates proper message handling with LangChain types.
    """
    llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
    
    # Get current messages from state
    messages = state.get("messages", [])
    
    # If this is the first iteration, add a system message
    if state.get("iteration_count", 0) == 0:
        system_message = SystemMessage(
            content="You are a knowledgeable assistant specializing in "
                    "explaining technical concepts clearly and concisely."
        )
        messages = [system_message] + list(messages)
    
    # Call the LLM with the properly formatted messages
    response = llm.invoke(messages)
    
    # The response is already an AIMessage object
    print(f"Assistant response: {response.content[:100]}...")
    
    return {
        "messages": [response],
        "iteration_count": state.get("iteration_count", 0) + 1
    }

This proper_chat_node shows best practices for working with LangChain messages. The LLM's invoke method accepts a list of BaseMessage objects and returns an AIMessage object, which we can directly add to our state.

Let's create a helper function to make it easy to add user messages:

def add_user_message(current_state: ProperChatState, user_input: str) -> ProperChatState:
    """
    Helper function to add a user message to the current state.
    """
    new_message = HumanMessage(content=user_input)
    
    # Create updated state with the new message
    updated_state = current_state.copy()
    updated_state["messages"] = list(current_state.get("messages", [])) + [new_message]
    
    return updated_state

Now let's build a complete conversational graph using proper message types:

from langgraph.graph import StateGraph, END

def create_proper_chat_graph():
    """
    Creates a chat graph using proper LangChain message types.
    """
    workflow = StateGraph(ProperChatState)
    
    workflow.add_node("chat", proper_chat_node)
    
    workflow.set_entry_point("chat")
    workflow.add_edge("chat", END)
    
    app = workflow.compile()
    return app

Let's use this graph in a multi-turn conversation:

app = create_proper_chat_graph()

# Initialize state with first user message
state = {
    "messages": [HumanMessage(content="What is LangGraph?")],
    "iteration_count": 0
}

# First turn
result = app.invoke(state)
print(f"Turn 1 - User: {result['messages'][0].content}")
print(f"Turn 1 - Assistant: {result['messages'][1].content[:200]}...")

# Add follow-up question
state = add_user_message(
    result,
    "Can you give me a simple example of how to use it?"
)

# Second turn
result = app.invoke(state)
print(f"\nTurn 2 - User: {result['messages'][2].content}")
print(f"Turn 2 - Assistant: {result['messages'][3].content[:200]}...")

This example demonstrates how to maintain a multi-turn conversation using proper message types. Each invocation of the graph builds upon the previous state, maintaining the full conversation history.

BUILDING A MULTI-AGENT RESEARCH SYSTEM: PRACTICAL APPLICATION

Now let's apply everything we've learned to build a practical multi-agent research system. This system will have multiple specialized agents that work together to research a topic, search for information, analyze findings, and synthesize a final answer.

First, let's define a comprehensive state for our research system:

from typing import TypedDict, Annotated, Sequence, Optional
from langchain.schema import BaseMessage
import operator

class ResearchAgentState(TypedDict):
    # The original research question
    question: str
    # Conversation messages between agents
    messages: Annotated[Sequence[BaseMessage], operator.add]
    # Search queries generated by the planner
    search_queries: Annotated[list[str], operator.add]
    # Results from web searches
    search_results: Annotated[list[dict], operator.add]
    # Analysis of the search results
    analysis: Annotated[list[str], operator.add]
    # The final synthesized answer
    final_answer: Optional[str]
    # Current step in the research process
    current_step: str
    # Number of iterations performed
    iteration_count: int
    # Maximum iterations allowed
    max_iterations: int

Now let's create the specialized agent nodes. First, a planner agent that generates search queries:

from langchain_openai import ChatOpenAI
from langchain.schema import SystemMessage, HumanMessage

def planner_agent_node(state: ResearchAgentState) -> dict:
    """
    The planner agent analyzes the research question and generates
    appropriate search queries to gather information.
    """
    llm = ChatOpenAI(model="gpt-4", temperature=0.3)
    
    question = state.get("question", "")
    existing_queries = state.get("search_queries", [])
    
    # Create a prompt for the planner
    system_prompt = SystemMessage(
        content="You are a research planner. Your job is to break down "
                "research questions into specific, targeted search queries. "
                "Generate 2-3 search queries that will help answer the question."
    )
    
    user_prompt = HumanMessage(
        content=f"Research question: {question}\n\n"
                f"Existing queries: {existing_queries}\n\n"
                f"Generate new search queries to gather comprehensive information."
    )
    
    response = llm.invoke([system_prompt, user_prompt])
    
    # Parse the response to extract queries (simplified for demonstration)
    # In a real system, you'd use structured output or parsing
    queries = [q.strip() for q in response.content.split("\n") if q.strip()]
    
    print(f"Planner generated {len(queries)} new queries")
    
    return {
        "search_queries": queries,
        "messages": [response],
        "current_step": "planning_complete"
    }

Next, a searcher agent that executes the search queries:

def searcher_agent_node(state: ResearchAgentState) -> dict:
    """
    The searcher agent executes search queries and retrieves results.
    In a real implementation, this would call actual search APIs.
    """
    queries = state.get("search_queries", [])
    
    print(f"Searcher executing {len(queries)} queries")
    
    # Simulate search results (in reality, call actual search API)
    all_results = []
    for query in queries:
        results = [
            {
                "query": query,
                "title": f"Result 1 for {query}",
                "snippet": f"This is detailed information about {query}. "
                          f"It contains relevant facts and data...",
                "url": f"https://example.com/{query.replace(' ', '-')}"
            },
            {
                "query": query,
                "title": f"Result 2 for {query}",
                "snippet": f"Additional information regarding {query}. "
                          f"This provides a different perspective...",
                "url": f"https://example.org/{query.replace(' ', '-')}"
            }
        ]
        all_results.extend(results)
    
    return {
        "search_results": all_results,
        "current_step": "search_complete"
    }

Now an analyzer agent that processes the search results:

def analyzer_agent_node(state: ResearchAgentState) -> dict:
    """
    The analyzer agent examines search results and extracts
    key information relevant to the research question.
    """
    llm = ChatOpenAI(model="gpt-4", temperature=0.3)
    
    question = state.get("question", "")
    results = state.get("search_results", [])
    
    # Create analysis prompt
    system_prompt = SystemMessage(
        content="You are a research analyst. Analyze search results and "
                "extract key information relevant to the research question. "
                "Be thorough and identify important facts, patterns, and insights."
    )
    
    # Format search results for analysis
    results_text = "\n\n".join([
        f"Source: {r['title']}\n{r['snippet']}"
        for r in results[-6:]  # Analyze last 6 results
    ])
    
    user_prompt = HumanMessage(
        content=f"Research question: {question}\n\n"
                f"Search results:\n{results_text}\n\n"
                f"Provide a detailed analysis of these results."
    )
    
    response = llm.invoke([system_prompt, user_prompt])
    
    print(f"Analyzer completed analysis: {response.content[:100]}...")
    
    return {
        "analysis": [response.content],
        "messages": [response],
        "current_step": "analysis_complete"
    }

Finally, a synthesizer agent that creates the final answer:

def synthesizer_agent_node(state: ResearchAgentState) -> dict:
    """
    The synthesizer agent combines all analyses into a
    comprehensive final answer to the research question.
    """
    llm = ChatOpenAI(model="gpt-4", temperature=0.5)
    
    question = state.get("question", "")
    analyses = state.get("analysis", [])
    
    system_prompt = SystemMessage(
        content="You are a research synthesizer. Your job is to combine "
                "multiple analyses into a clear, comprehensive answer. "
                "Provide a well-structured response that directly addresses "
                "the research question."
    )
    
    # Combine all analyses
    combined_analysis = "\n\n".join([
        f"Analysis {i+1}:\n{analysis}"
        for i, analysis in enumerate(analyses)
    ])
    
    user_prompt = HumanMessage(
        content=f"Research question: {question}\n\n"
                f"Analyses:\n{combined_analysis}\n\n"
                f"Synthesize a comprehensive final answer."
    )
    
    response = llm.invoke([system_prompt, user_prompt])
    
    print(f"Synthesizer created final answer: {response.content[:100]}...")
    
    return {
        "final_answer": response.content,
        "messages": [response],
        "current_step": "synthesis_complete"
    }

Now we need a router function to coordinate these agents:

from typing import Literal

def research_router(
    state: ResearchAgentState
) -> Literal["planner", "searcher", "analyzer", "synthesizer", "end"]:
    """
    Routes the workflow between different research agents based
    on the current step and iteration count.
    """
    current_step = state.get("current_step", "start")
    iteration = state.get("iteration_count", 0)
    max_iterations = state.get("max_iterations", 2)
    
    # Check if we've reached maximum iterations
    if iteration >= max_iterations:
        # If we have analysis, synthesize; otherwise end
        if state.get("analysis"):
            if current_step != "synthesis_complete":
                return "synthesizer"
        return "end"
    
    # Route based on current step
    if current_step == "start":
        return "planner"
    elif current_step == "planning_complete":
        return "searcher"
    elif current_step == "search_complete":
        return "analyzer"
    elif current_step == "analysis_complete":
        # Decide whether to iterate or synthesize
        if iteration < max_iterations - 1:
            return "planner"  # Do another iteration
        else:
            return "synthesizer"
    elif current_step == "synthesis_complete":
        return "end"
    
    return "end"

Now let's build the complete multi-agent research graph:

from langgraph.graph import StateGraph, END

def create_research_agent_graph():
    """
    Creates a multi-agent research system with planner, searcher,
    analyzer, and synthesizer agents working together.
    """
    workflow = StateGraph(ResearchAgentState)
    
    # Add all agent nodes
    workflow.add_node("planner", planner_agent_node)
    workflow.add_node("searcher", searcher_agent_node)
    workflow.add_node("analyzer", analyzer_agent_node)
    workflow.add_node("synthesizer", synthesizer_agent_node)
    
    # Set entry point
    workflow.set_entry_point("planner")
    
    # Add conditional edges from each node using the router
    for node_name in ["planner", "searcher", "analyzer", "synthesizer"]:
        workflow.add_conditional_edges(
            node_name,
            research_router,
            {
                "planner": "planner",
                "searcher": "searcher",
                "analyzer": "analyzer",
                "synthesizer": "synthesizer",
                "end": END
            }
        )
    
    # Compile the graph
    app = workflow.compile()
    return app

Let's use our multi-agent research system:

# Create the research agent graph
research_app = create_research_agent_graph()

# Define a research question
initial_state = {
    "question": "What are the key benefits and challenges of using "
               "LangGraph for building multi-agent LLM applications?",
    "messages": [],
    "search_queries": [],
    "search_results": [],
    "analysis": [],
    "final_answer": None,
    "current_step": "start",
    "iteration_count": 0,
    "max_iterations": 2
}

# Execute the research workflow
final_state = research_app.invoke(initial_state)

# Display results
print("\n" + "="*80)
print("RESEARCH COMPLETE")
print("="*80)
print(f"\nQuestion: {final_state['question']}")
print(f"\nQueries executed: {len(final_state['search_queries'])}")
print(f"Results gathered: {len(final_state['search_results'])}")
print(f"Iterations: {final_state['iteration_count']}")
print(f"\nFinal Answer:\n{final_state['final_answer']}")

This multi-agent research system demonstrates the power of LangGraph. Multiple specialized agents work together, each handling a specific aspect of the research process. The router coordinates their activities, and the shared state allows them to build upon each other's work.

PERSISTENCE AND CHECKPOINTING: SAVING YOUR WORKFLOW STATE

One of LangGraph's powerful features is the ability to persist workflow state and create checkpoints. This allows you to pause and resume workflows, implement human-in-the-loop patterns, and recover from failures.

To enable persistence, you need to provide a checkpointer when compiling your graph. LangGraph supports various checkpointer implementations. Let's use the MemorySaver for demonstration:

from langgraph.checkpoint.memory import MemorySaver

def create_persistent_chat_graph():
    """
    Creates a chat graph with state persistence enabled.
    """
    workflow = StateGraph(ProperChatState)
    
    workflow.add_node("chat", proper_chat_node)
    workflow.set_entry_point("chat")
    workflow.add_edge("chat", END)
    
    # Create a memory-based checkpointer
    memory = MemorySaver()
    
    # Compile with checkpointer
    app = workflow.compile(checkpointer=memory)
    
    return app

When using a persistent graph, you need to provide a thread_id to identify different conversation threads:

persistent_app = create_persistent_chat_graph()

# Configuration with thread ID
config = {"configurable": {"thread_id": "conversation-1"}}

# First message in thread
state1 = {
    "messages": [HumanMessage(content="Hello, what is LangGraph?")],
    "iteration_count": 0
}

result1 = persistent_app.invoke(state1, config)
print(f"Response 1: {result1['messages'][-1].content[:100]}...")

# Continue the same thread with a follow-up
state2 = {
    "messages": [HumanMessage(content="Can you give me an example?")],
    "iteration_count": result1["iteration_count"]
}

result2 = persistent_app.invoke(state2, config)
print(f"Response 2: {result2['messages'][-1].content[:100]}...")

# The checkpointer maintains the full conversation history
print(f"\nTotal messages in thread: {len(result2['messages'])}")

The checkpointer automatically saves the state after each node execution. This enables powerful patterns like human-in-the-loop workflows where you can pause execution, get human input, and then resume.

HUMAN-IN-THE-LOOP PATTERNS: INTERACTIVE WORKFLOWS

LangGraph makes it easy to implement human-in-the-loop patterns where human input is required at certain points in the workflow. Let's build an example where a human reviewer approves or rejects content before it's finalized.

First, let's define a state that tracks approval status:

from typing import TypedDict, Annotated, Sequence, Optional, Literal
from langchain.schema import BaseMessage
import operator

class ApprovalState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    draft_content: Optional[str]
    human_feedback: Optional[str]
    approval_status: Optional[Literal["pending", "approved", "rejected"]]
    final_content: Optional[str]

Now let's create nodes for content generation and revision:

from langchain_openai import ChatOpenAI
from langchain.schema import SystemMessage, HumanMessage

def generate_content_node(state: ApprovalState) -> dict:
    """
    Generates initial draft content based on the user's request.
    """
    llm = ChatOpenAI(model="gpt-4", temperature=0.7)
    
    messages = state.get("messages", [])
    
    system_prompt = SystemMessage(
        content="You are a content writer. Create high-quality content "
                "based on the user's request."
    )
    
    response = llm.invoke([system_prompt] + list(messages))
    
    print(f"Generated draft content: {response.content[:100]}...")
    
    return {
        "draft_content": response.content,
        "messages": [response],
        "approval_status": "pending"
    }

def revise_content_node(state: ApprovalState) -> dict:
    """
    Revises content based on human feedback.
    """
    llm = ChatOpenAI(model="gpt-4", temperature=0.7)
    
    draft = state.get("draft_content", "")
    feedback = state.get("human_feedback", "")
    
    system_prompt = SystemMessage(
        content="You are a content editor. Revise the draft based on "
                "the feedback provided."
    )
    
    user_prompt = HumanMessage(
        content=f"Draft:\n{draft}\n\nFeedback:\n{feedback}\n\n"
                f"Please revise the content accordingly."
    )
    
    response = llm.invoke([system_prompt, user_prompt])
    
    print(f"Revised content: {response.content[:100]}...")
    
    return {
        "draft_content": response.content,
        "messages": [response],
        "approval_status": "pending"
    }

def finalize_content_node(state: ApprovalState) -> dict:
    """
    Finalizes approved content.
    """
    draft = state.get("draft_content", "")
    
    print("Content approved and finalized!")
    
    return {
        "final_content": draft,
        "approval_status": "approved"
    }

Now let's create a router that handles the approval workflow:

def approval_router(
    state: ApprovalState
) -> Literal["generate", "revise", "finalize", "human_review"]:
    """
    Routes based on approval status and human feedback.
    """
    status = state.get("approval_status")
    
    if status is None:
        return "generate"
    elif status == "pending":
        return "human_review"
    elif status == "rejected":
        return "revise"
    elif status == "approved":
        return "finalize"
    
    return "human_review"

Here's how you would build the graph with human-in-the-loop:

from langgraph.graph import StateGraph, END

def create_approval_workflow():
    """
    Creates a workflow that requires human approval.
    """
    workflow = StateGraph(ApprovalState)
    
    workflow.add_node("generate", generate_content_node)
    workflow.add_node("revise", revise_content_node)
    workflow.add_node("finalize", finalize_content_node)
    
    workflow.set_entry_point("generate")
    
    # After generation, always go to human review (simulated)
    workflow.add_edge("generate", "finalize")
    workflow.add_edge("revise", "finalize")
    workflow.add_edge("finalize", END)
    
    app = workflow.compile()
    return app

In a real implementation with checkpointing, you would pause execution before the human review step, wait for human input, and then resume with the updated state.

STREAMING OUTPUTS: REAL-TIME FEEDBACK

LangGraph supports streaming, which allows you to get real-time updates as your graph executes. This is particularly useful for long-running workflows or when you want to provide immediate feedback to users.

Here's how to use streaming:

# Create a graph (using our research agent as an example)
research_app = create_research_agent_graph()

initial_state = {
    "question": "What is LangGraph?",
    "messages": [],
    "search_queries": [],
    "search_results": [],
    "analysis": [],
    "final_answer": None,
    "current_step": "start",
    "iteration_count": 0,
    "max_iterations": 1
}

# Stream the execution
print("Streaming research workflow:")
print("-" * 80)

for output in research_app.stream(initial_state):
    # Each output is a dictionary with node name as key
    for node_name, node_output in output.items():
        print(f"\nNode '{node_name}' completed")
        print(f"Current step: {node_output.get('current_step', 'N/A')}")
        
        # You can access any part of the state here
        if 'final_answer' in node_output and node_output['final_answer']:
            print(f"Final answer ready: {node_output['final_answer'][:100]}...")

print("\n" + "-" * 80)
print("Workflow complete!")

The stream method yields the output of each node as it completes, allowing you to provide real-time progress updates to users or log detailed execution information.

BEST PRACTICES AND PATTERNS

As you build more complex LangGraph applications, following these best practices will help you create maintainable, efficient, and reliable systems.

First, design your state schema carefully. Your state should contain all the information that needs to flow between nodes, but avoid making it overly complex. Group related information together and use clear, descriptive field names. Use Optional types for fields that might not always be present.

Second, keep your nodes focused and single-purpose. Each node should perform one clear task. This makes your graph easier to understand, test, and debug. If a node is doing too many things, consider splitting it into multiple nodes.

Third, use meaningful node names that clearly describe what the node does. Names like "planner", "searcher", and "analyzer" are much better than "node1", "node2", and "node3". Good names make your graph self-documenting.

Fourth, implement proper error handling in your nodes. Wrap LLM calls and external API calls in try-except blocks. When an error occurs, update the state to reflect the error condition so your router can handle it appropriately.

Here's an example of a node with proper error handling:

def robust_llm_node(state: ProperChatState) -> dict:
    """
    An LLM node with comprehensive error handling.
    """
    try:
        llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
        messages = state.get("messages", [])
        
        if not messages:
            raise ValueError("No messages to process")
        
        response = llm.invoke(messages)
        
        return {
            "messages": [response],
            "iteration_count": state.get("iteration_count", 0) + 1
        }
        
    except Exception as e:
        print(f"Error in LLM node: {str(e)}")
        
        # Return an error message in the state
        error_message = AIMessage(
            content=f"I encountered an error: {str(e)}. "
                   f"Please try rephrasing your question."
        )
        
        return {
            "messages": [error_message],
            "iteration_count": state.get("iteration_count", 0) + 1
        }

Fifth, use type hints consistently throughout your code. This helps catch errors early and makes your code more maintainable. LangGraph works well with Python's type system, so take advantage of it.

Sixth, test your nodes independently before integrating them into a graph. Each node is just a Python function, so you can easily write unit tests for them:

def test_increment_counter():
    """
    Test the increment counter node.
    """
    test_state = {"counter": 5, "message": "test"}
    result = increment_counter_node(test_state)
    
    assert result["counter"] == 6
    print("Test passed: Counter incremented correctly")

test_increment_counter()

Seventh, use logging to track execution flow. This is invaluable for debugging complex graphs:

import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def logged_node(state: SimpleState) -> dict:
    """
    A node that logs its execution.
    """
    logger.info(f"Node executing with state: {state}")
    
    result = {"counter": state.get("counter", 0) + 1}
    
    logger.info(f"Node returning: {result}")
    
    return result

Eighth, when building multi-agent systems, clearly define each agent's responsibility. Avoid overlap between agents. Each agent should have a distinct role that contributes to the overall goal.

Ninth, use conditional edges to implement retry logic and error recovery. If a node fails or produces unsatisfactory results, your router can direct execution to a retry node or an alternative path.

Tenth, for production systems, use persistent checkpointers (not just MemorySaver) to ensure state is preserved across application restarts. LangGraph supports various backend storage options for checkpointing.

ADVANCED PATTERNS: SUBGRAPHS AND COMPOSITION

As your applications grow more complex, you may want to compose multiple graphs together. LangGraph supports this through subgraphs, where one graph can be used as a node in another graph.

Let's create a simple example with a subgraph:

from langgraph.graph import StateGraph, END

# Define state for the subgraph
class SubGraphState(TypedDict):
    input_value: int
    output_value: int

def double_node(state: SubGraphState) -> dict:
    """
    Doubles the input value.
    """
    value = state.get("input_value", 0)
    return {"output_value": value * 2}

def create_doubling_subgraph():
    """
    Creates a simple subgraph that doubles a value.
    """
    workflow = StateGraph(SubGraphState)
    workflow.add_node("double", double_node)
    workflow.set_entry_point("double")
    workflow.add_edge("double", END)
    
    return workflow.compile()

Now you can use this subgraph as a node in a larger graph. This pattern is useful for organizing complex workflows into modular, reusable components.

CONCLUSION: YOUR JOURNEY WITH LANGGRAPH

Congratulations! You have now learned the fundamental concepts and patterns for building sophisticated multi-agent LLM applications with LangGraph. Let's recap what we've covered.

We started by understanding what LangGraph is and why it's valuable for building complex LLM applications. We learned that LangGraph provides a graph-based framework for orchestrating multiple LLM calls, managing state, and implementing conditional logic.

We explored the four core concepts: State, which represents the shared information flowing through your application; Graphs, which define the overall structure; Nodes, which perform discrete units of work; and Edges, which control execution flow.

We learned how to define state schemas using TypedDict and how to use the Annotated type with operator.add to accumulate values rather than replace them. This is crucial for maintaining conversation history and building up context.

We created various types of nodes, from simple functions that increment counters to sophisticated agents that call LLMs, perform web searches, analyze data, and synthesize results. We saw how nodes receive state, perform operations, and return state updates.

We explored both normal edges for sequential flow and conditional edges for decision-making. We learned how to write router functions that examine state and determine the next node to execute, enabling dynamic, intelligent workflows.

We built a complete multi-agent research system with specialized agents for planning, searching, analyzing, and synthesizing information. This demonstrated how multiple agents can work together, coordinated by a router, to accomplish complex tasks.

We learned about persistence and checkpointing, which allow you to save workflow state, implement human-in-the-loop patterns, and recover from failures. We saw how to use streaming to get real-time updates as graphs execute.

Finally, we covered best practices including careful state design, focused single-purpose nodes, meaningful naming, error handling, type hints, testing, logging, and modular composition.

You now have the knowledge to build your own LangGraph applications. Start with simple graphs to get comfortable with the concepts, then gradually increase complexity as you gain confidence. Remember that the key to success with LangGraph is thinking in terms of graphs: what are the steps in your workflow, what information needs to flow between them, and what decisions need to be made along the way.

The LangGraph library continues to evolve with new features and capabilities. The patterns and concepts you've learned here provide a solid foundation that will serve you well as you explore more advanced features and build increasingly sophisticated applications.

Happy building, and may your LLM applications be stateful, intelligent, and powerful!

Friday, October 02, 2026

MASTERING LANGCHAIN: FROM SIMPLE CHAINS TO INTELLIGENT AGENTS

 



THE ORCHESTRATION LAYER FOR AI APPLICATIONS

Imagine you're building a house. You could craft every nail, cut every board, and mix every batch of concrete yourself. Or you could use pre-made materials and focus on the architecture. LangChain is the latter approach for AI applications.

When you built the HuggingFace chatbot in the previous tutorial, you manually connected components: loading models, creating embeddings, managing vector stores, formatting prompts, and orchestrating retrieval. It worked, but it required understanding and implementing every detail. LangChain provides a higher-level abstraction layer that handles these patterns for you.

Founded by Harrison Chase in late 2022, LangChain exploded in popularity because it solved a real problem: building LLM applications involves repetitive patterns. Every RAG system needs document loading, text splitting, embedding, vector storage, and retrieval. Every chatbot needs memory management. Every agent needs tool integration. LangChain provides battle-tested implementations of these patterns.

But LangChain is more than a convenience library. It introduces powerful abstractions that change how you think about AI applications. Instead of writing procedural code that calls APIs, you compose declarative chains that express what you want to happen. Instead of managing state manually, you use memory systems. Instead of hard-coding logic, you create agents that reason about which tools to use.

By the end of this tutorial, you'll understand LangChain's core concepts and abstractions. You'll build chatbots with memory, implement RAG systems with just a few lines of code, and create agents that can use tools autonomously. You'll understand when to use LangChain and when lower-level libraries like HuggingFace are more appropriate.

Let's dive into the world of LangChain and discover how it transforms AI application development.

UNDERSTANDING LANGCHAIN'S PHILOSOPHY

Before writing code, let's understand LangChain's design philosophy. This will help you think in "LangChain terms" and use the library effectively.

LangChain is built around several core principles. First is composability. Complex applications are built by composing simple components. A RAG system composes a retriever with a language model. An agent composes tools with a reasoning engine. This compositional approach makes systems easier to understand, test, and modify.

Second is abstraction. LangChain provides abstract interfaces for common components. Whether you use OpenAI, Anthropic, or a local model, the interface remains the same. Whether you store vectors in FAISS, Pinecone, or Chroma, the retriever interface is consistent. This abstraction lets you swap components without rewriting your application.

Third is standardization. LangChain establishes patterns for common tasks. There's a standard way to handle chat history, a standard way to structure prompts, a standard way to implement retrieval. These patterns emerge from real-world usage and represent best practices.

Fourth is extensibility. While LangChain provides many built-in components, you can easily create custom ones. Need a special document loader? Implement the base class. Need a custom tool for your agent? Define its interface. The framework is designed for extension.

Understanding these principles helps you use LangChain effectively. You're not just calling functions; you're composing components into systems.

SETTING UP YOUR LANGCHAIN ENVIRONMENT

Let's prepare our development environment. LangChain has a modular structure, so we'll install the components we need.

pip install langchain langchain-community langchain-openai
pip install faiss-cpu sentence-transformers
pip install python-dotenv

The langchain package contains the core abstractions and base classes. The langchain-community package includes integrations with various services and tools. The langchain-openai package provides OpenAI integrations. We'll also install FAISS for vector storage and sentence-transformers for embeddings.

LangChain works with many LLM providers. For this tutorial, we'll use OpenAI's API because it's reliable and well-documented. You'll need an API key from OpenAI. Create a file named dot-env in your project directory:

OPENAI_API_KEY=your-api-key-here

Now let's verify the installation:

import langchain
from langchain_openai import ChatOpenAI
from dotenv import load_dotenv
import os

# Load environment variables
load_dotenv()

# Verify API key is loaded
api_key = os.getenv("OPENAI_API_KEY")
if api_key:
    print("API key loaded successfully")
    print(f"LangChain version: {langchain.__version__}")
else:
    print("Warning: OPENAI_API_KEY not found in environment")

If you prefer to use local models instead of OpenAI, you can use Ollama or HuggingFace models. We'll show alternatives throughout the tutorial.

YOUR FIRST LANGCHAIN INTERACTION

Let's start with the simplest possible LangChain program: asking a question to an LLM.

from langchain_openai import ChatOpenAI
from dotenv import load_dotenv

# Load environment variables
load_dotenv()

# Create a language model instance
llm = ChatOpenAI(
    model="gpt-3.5-turbo",
    temperature=0.7
)

# Ask a question
response = llm.invoke("What are the three laws of robotics?")
print(response.content)

This code creates a ChatOpenAI instance, which is LangChain's wrapper around OpenAI's chat models. The invoke method sends a message and returns the response. The temperature parameter controls randomness, just like in the HuggingFace tutorial.

The response object contains more than just text. Let's explore it:

response = llm.invoke("Explain quantum computing in one sentence")

print(f"Content: {response.content}")
print(f"Response metadata: {response.response_metadata}")
print(f"Type: {type(response)}")

The response is an AIMessage object containing the content, metadata about the API call (like token usage), and other information. This structured response makes it easy to extract what you need.

Now let's see LangChain's real power: handling conversations with multiple messages.

from langchain_core.messages import HumanMessage, SystemMessage, AIMessage

# Create a conversation
messages = [
    SystemMessage(content="You are a helpful physics tutor."),
    HumanMessage(content="What is quantum entanglement?"),
]

response = llm.invoke(messages)
print(f"Assistant: {response.content}\n")

# Continue the conversation
messages.append(response)
messages.append(HumanMessage(content="Can you give me a simple analogy?"))

response = llm.invoke(messages)
print(f"Assistant: {response.content}")

LangChain uses message objects to represent different roles in a conversation. SystemMessage sets the AI's behavior and context. HumanMessage represents user input. AIMessage represents the AI's responses. This structure mirrors how chat models actually work.

The conversation maintains context because we pass the entire message history with each call. The model sees the previous exchange and can provide coherent follow-up responses.

PROMPT TEMPLATES: STRUCTURED PROMPT ENGINEERING

Hard-coding prompts works for simple cases, but real applications need dynamic prompts that incorporate variables. LangChain's prompt templates solve this elegantly.

from langchain_core.prompts import ChatPromptTemplate

# Create a prompt template
template = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant that translates {input_language} to {output_language}."),
    ("human", "{text}")
])

# Format the prompt with variables
messages = template.format_messages(
    input_language="English",
    output_language="French",
    text="Hello, how are you?"
)

# Use with the LLM
response = llm.invoke(messages)
print(response.content)

The template uses curly braces for variables. When you call format_messages, it substitutes the variables with actual values. This separation of template and data makes prompts reusable and testable.

You can create more complex templates with multiple variables and conditional logic:

from langchain_core.prompts import PromptTemplate

# Template for code explanation
code_template = PromptTemplate(
    input_variables=["language", "code", "detail_level"],
    template="""You are an expert {language} programmer. 

Explain the following code at a {detail_level} level of detail.

Code: {code}

Explanation:""" )

# Use the template
prompt = code_template.format(
    language="Python",
    detail_level="beginner-friendly",
    code="def fibonacci(n):\n    return n if n <= 1 else fibonacci(n-1) + fibonacci(n-2)"
)

response = llm.invoke(prompt)
print(response.content)

Templates support different formats for different use cases. ChatPromptTemplate is for chat models, while PromptTemplate is for completion models. There are also specialized templates for few-shot learning and other patterns.

Here's a practical example showing few-shot prompting:

from langchain_core.prompts import FewShotPromptTemplate, PromptTemplate

# Examples for few-shot learning
examples = [
    {
        "input": "The movie was fantastic!",
        "output": "Positive"
    },
    {
        "input": "I hated every minute of it.",
        "output": "Negative"
    },
    {
        "input": "It was okay, nothing special.",
        "output": "Neutral"
    }
]

# Template for each example
example_template = PromptTemplate(
    input_variables=["input", "output"],
    template="Input: {input}\nSentiment: {output}"
)

# Few-shot prompt template
few_shot_template = FewShotPromptTemplate(
    examples=examples,
    example_prompt=example_template,
    prefix="Classify the sentiment of the following text.",
    suffix="Input: {input}\nSentiment:",
    input_variables=["input"]
)

# Use the template
prompt = few_shot_template.format(input="This product exceeded my expectations!")
response = llm.invoke(prompt)
print(response.content)

Few-shot prompting provides examples to guide the model's behavior. This is especially useful for tasks where you want consistent output formatting or specific classification categories.

CHAINS: COMPOSING OPERATIONS

Now we reach one of LangChain's most powerful concepts: chains. A chain is a sequence of operations that process data. The output of one step becomes the input to the next.

The simplest chain connects a prompt template to an LLM:

from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI

# Create components
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a creative storyteller."),
    ("human", "Write a one-paragraph story about {topic}")
])

# Create a chain using the pipe operator
chain = prompt | llm

# Invoke the chain
response = chain.invoke({"topic": "a robot learning to paint"})
print(response.content)

The pipe operator (vertical bar) creates a chain. When you invoke the chain with a dictionary of variables, it flows through each component. The prompt template formats the messages, then the LLM generates a response.

You can extend chains with output parsers to structure the response:

from langchain_core.output_parsers import StrOutputParser

# Add an output parser to extract just the string content
chain = prompt | llm | StrOutputParser()

# Now the result is a string, not an AIMessage object
result = chain.invoke({"topic": "a time-traveling historian"})
print(f"Type: {type(result)}")
print(f"Story: {result}")

The StrOutputParser extracts the content string from the AIMessage. This makes the chain's output cleaner and easier to use in subsequent operations.

Let's build a more complex chain that performs multiple steps:

from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain_core.output_parsers import StrOutputParser

# First chain: Generate a topic
topic_prompt = ChatPromptTemplate.from_template(
    "Suggest an interesting topic for a {genre} story. Just give the topic, nothing else."
)
topic_chain = topic_prompt | llm | StrOutputParser()

# Second chain: Write the story
story_prompt = ChatPromptTemplate.from_template(
    "Write a short {genre} story about: {topic}"
)
story_chain = story_prompt | llm | StrOutputParser()

# Combine the chains
def generate_story(genre):
    """Generate a story by first creating a topic, then writing about it."""
    # Get a topic
    topic = topic_chain.invoke({"genre": genre})
    print(f"Generated topic: {topic}\n")
    
    # Write the story
    story = story_chain.invoke({"genre": genre, "topic": topic})
    return story

# Use the combined workflow
result = generate_story("science fiction")
print(result)

This demonstrates sequential processing where the output of one chain feeds into another. The first chain generates a topic, and the second chain uses that topic to write a story.

LangChain provides specialized chain types for common patterns. The LLMChain is a simple prompt-to-LLM chain:

from langchain.chains import LLMChain

# Create an LLMChain
prompt = ChatPromptTemplate.from_template("What is the capital of {country}?")
chain = LLMChain(llm=llm, prompt=prompt)

# Use it
result = chain.invoke({"country": "Japan"})
print(result["text"])

While the modern approach uses the pipe operator, LLMChain is still useful for backward compatibility and certain use cases.

MEMORY: MAINTAINING CONVERSATION CONTEXT

Real chatbots need to remember previous exchanges. LangChain's memory systems handle this automatically.

from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationChain

# Create a memory instance
memory = ConversationBufferMemory()

# Create a conversation chain with memory
conversation = ConversationChain(
    llm=llm,
    memory=memory,
    verbose=True  # Shows what's happening internally
)

# Have a conversation
print(conversation.predict(input="Hi, my name is Alice"))
print("\n" + "="*50 + "\n")

print(conversation.predict(input="What's 2+2?"))
print("\n" + "="*50 + "\n")

print(conversation.predict(input="What's my name?"))

The ConversationBufferMemory stores all messages in a buffer. When you ask "What's my name?", the model can answer because it has access to the entire conversation history.

The verbose flag shows you what's being sent to the LLM. You'll see that each call includes the full conversation history, allowing the model to maintain context.

Let's examine the memory directly:

# View the conversation history
print("\nConversation history:")
print(memory.load_memory_variables({}))

# Clear the memory
memory.clear()
print("\nMemory cleared")

The load_memory_variables method returns the stored conversation. This is useful for debugging or saving conversations.

LangChain offers different memory types for different needs. ConversationBufferWindowMemory keeps only the last N messages:

from langchain.memory import ConversationBufferWindowMemory

# Keep only the last 3 exchanges
windowed_memory = ConversationBufferWindowMemory(k=3)

conversation = ConversationChain(
    llm=llm,
    memory=windowed_memory
)

# Have a longer conversation
conversation.predict(input="My favorite color is blue")
conversation.predict(input="I work as a software engineer")
conversation.predict(input="I enjoy hiking on weekends")
conversation.predict(input="I have a cat named Whiskers")

# This will only remember the last 3 exchanges
response = conversation.predict(input="What's my favorite color?")
print(response)

The model might not remember your favorite color because that exchange has been pushed out of the window. This memory type is useful for long conversations where you want to limit token usage.

Another useful memory type is ConversationSummaryMemory, which summarizes old messages:

from langchain.memory import ConversationSummaryMemory

# Create summary memory
summary_memory = ConversationSummaryMemory(llm=llm)

conversation = ConversationChain(
    llm=llm,
    memory=summary_memory,
    verbose=True
)

# Have a conversation
conversation.predict(input="I'm planning a trip to Japan next spring")
conversation.predict(input="I want to visit Tokyo, Kyoto, and Osaka")
conversation.predict(input="I'm particularly interested in traditional temples")

# Check the summary
print("\nMemory summary:")
print(summary_memory.load_memory_variables({}))

The summary memory uses the LLM to create a running summary of the conversation. This keeps token usage low while maintaining important context.

For more control, you can use ConversationBufferMemory with custom keys:

from langchain.memory import ConversationBufferMemory
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder

# Create memory with custom key
memory = ConversationBufferMemory(
    memory_key="chat_history",
    return_messages=True
)

# Create a prompt that uses the memory
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant."),
    MessagesPlaceholder(variable_name="chat_history"),
    ("human", "{input}")
])

# Create chain with memory
from langchain.chains import LLMChain

chain = LLMChain(
    llm=llm,
    prompt=prompt,
    memory=memory
)

# Use the chain
response = chain.predict(input="Hi, I'm learning LangChain")
print(response)

response = chain.predict(input="What am I learning?")
print(response)

The MessagesPlaceholder reserves a spot in the prompt for the conversation history. This gives you fine-grained control over where history appears in your prompts.

DOCUMENT LOADERS: INGESTING INFORMATION

Now we're ready to build RAG systems with LangChain. The first step is loading documents. LangChain provides loaders for many file types.

from langchain_community.document_loaders import TextLoader

# Load a text file
loader = TextLoader("example.txt")
documents = loader.load()

print(f"Loaded {len(documents)} document(s)")
print(f"First document preview: {documents[0].page_content[:200]}")
print(f"Metadata: {documents[0].metadata}")

Each document has page_content (the text) and metadata (information about the source). The metadata typically includes the filename and other relevant details.

For this tutorial, let's create documents programmatically:

from langchain_core.documents import Document

# Create sample documents about LangChain
documents = [
    Document(
        page_content="""LangChain is a framework for developing applications powered by 
        language models. It was created by Harrison Chase and released in October 2022. 
        The framework provides abstractions for working with LLMs, including chains, 
        agents, and memory systems.""",
        metadata={"source": "intro", "topic": "overview"}
    ),
    Document(
        page_content="""Chains in LangChain are sequences of operations that process data. 
        The simplest chain connects a prompt template to an LLM. More complex chains can 
        include multiple steps, conditional logic, and parallel execution. Chains are 
        composable, meaning you can combine simple chains into complex workflows.""",
        metadata={"source": "chains", "topic": "concepts"}
    ),
    Document(
        page_content="""Memory systems in LangChain maintain conversation context. 
        ConversationBufferMemory stores all messages. ConversationBufferWindowMemory 
        keeps only recent messages. ConversationSummaryMemory creates summaries of old 
        messages. Each memory type offers different tradeoffs between context preservation 
        and token usage.""",
        metadata={"source": "memory", "topic": "concepts"}
    ),
    Document(
        page_content="""Agents in LangChain can use tools to accomplish tasks. An agent 
        receives a task, reasons about which tools to use, executes the tools, and 
        synthesizes the results. Tools can be anything from calculators to search engines 
        to custom APIs. Agents enable autonomous behavior where the LLM decides what 
        actions to take.""",
        metadata={"source": "agents", "topic": "advanced"}
    )
]

print(f"Created {len(documents)} documents")

LangChain supports many document loaders. Here are some examples:

# PDF loader (requires pypdf)
# from langchain_community.document_loaders import PyPDFLoader
# loader = PyPDFLoader("document.pdf")
# documents = loader.load()

# Web page loader (requires beautifulsoup4)
# from langchain_community.document_loaders import WebBaseLoader
# loader = WebBaseLoader("https://example.com")
# documents = loader.load()

# CSV loader
# from langchain_community.document_loaders import CSVLoader
# loader = CSVLoader("data.csv")
# documents = loader.load()

# Directory loader (loads all files in a directory)
# from langchain_community.document_loaders import DirectoryLoader
# loader = DirectoryLoader("./documents", glob="**/*.txt")
# documents = loader.load()

Each loader handles the specifics of its file format, returning a consistent Document structure.

TEXT SPLITTERS: CHUNKING DOCUMENTS

Large documents need to be split into smaller chunks for effective retrieval. LangChain provides sophisticated text splitters.

from langchain.text_splitter import RecursiveCharacterTextSplitter

# Create a text splitter
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=200,
    chunk_overlap=50,
    length_function=len,
    separators=["\n\n", "\n", " ", ""]
)

# Split the documents
split_docs = text_splitter.split_documents(documents)

print(f"Original documents: {len(documents)}")
print(f"Split chunks: {len(split_docs)}")
print(f"\nFirst chunk:")
print(split_docs[0].page_content)
print(f"Metadata: {split_docs[0].metadata}")

The RecursiveCharacterTextSplitter tries to split on paragraph boundaries first, then sentences, then words, then characters. This preserves semantic coherence better than naive splitting.

The chunk_overlap parameter ensures that context at chunk boundaries isn't lost. If a sentence is split across chunks, the overlap captures it in both chunks.

Let's see how different splitters work:

from langchain.text_splitter import CharacterTextSplitter

# Simple character splitter
simple_splitter = CharacterTextSplitter(
    chunk_size=200,
    chunk_overlap=0,
    separator=" "
)

simple_chunks = simple_splitter.split_documents(documents)

print(f"Recursive splitter: {len(split_docs)} chunks")
print(f"Simple splitter: {len(simple_chunks)} chunks")

# Compare first chunks
print(f"\nRecursive first chunk:\n{split_docs[0].page_content}\n")
print(f"Simple first chunk:\n{simple_chunks[0].page_content}")

The recursive splitter generally produces more coherent chunks because it respects document structure.

For code, there's a specialized splitter:

from langchain.text_splitter import Language, RecursiveCharacterTextSplitter

# Python code splitter
python_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.PYTHON,
    chunk_size=500,
    chunk_overlap=50
)

python_code = """

def fibonacci(n): '''Calculate the nth Fibonacci number.''' if n <= 1: return n return fibonacci(n-1) + fibonacci(n-2)

class Calculator: '''A simple calculator class.'''

def add(self, a, b):
    return a + b

def multiply(self, a, b):
    return a * b

"""

code_chunks = python_splitter.split_text(python_code)
for i, chunk in enumerate(code_chunks):
    print(f"Chunk {i+1}:\n{chunk}\n")

The code splitter understands programming language structure and tries to keep functions and classes together.

VECTOR STORES: STORING AND SEARCHING EMBEDDINGS

Now we need to convert our chunks into embeddings and store them for retrieval. LangChain abstracts the vector store interface.

from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS

# Create an embedding model
embeddings = HuggingFaceEmbeddings(
    model_name="all-MiniLM-L6-v2"
)

# Create a vector store from documents
vectorstore = FAISS.from_documents(
    documents=split_docs,
    embedding=embeddings
)

print("Vector store created")
print(f"Number of vectors: {vectorstore.index.ntotal}")

The from_documents method handles everything: generating embeddings for each chunk and adding them to the FAISS index. The result is a searchable vector store.

Let's search the vector store:

# Search for similar documents
query = "What are chains in LangChain?"
results = vectorstore.similarity_search(query, k=2)

print(f"Query: {query}\n")
for i, doc in enumerate(results, 1):
    print(f"Result {i}:")
    print(f"Content: {doc.page_content}")
    print(f"Metadata: {doc.metadata}\n")

The similarity_search method finds the most relevant chunks. It returns Document objects with both content and metadata.

You can also get similarity scores:

# Search with scores
results_with_scores = vectorstore.similarity_search_with_score(query, k=3)

print(f"Query: {query}\n")
for doc, score in results_with_scores:
    print(f"Score: {score:.4f}")
    print(f"Content: {doc.page_content[:100]}...\n")

Lower scores indicate higher similarity (because FAISS uses L2 distance).

The vector store can be saved and loaded:

# Save the vector store
vectorstore.save_local("langchain_vectorstore")

# Load it later
loaded_vectorstore = FAISS.load_local(
    "langchain_vectorstore",
    embeddings,
    allow_dangerous_deserialization=True
)

print("Vector store loaded successfully")

This allows you to build the index once and reuse it across sessions.

LangChain supports many vector store backends:

# Chroma (requires chromadb)
# from langchain_community.vectorstores import Chroma
# vectorstore = Chroma.from_documents(documents=split_docs, embedding=embeddings)

# Pinecone (requires pinecone-client and API key)
# from langchain_community.vectorstores import Pinecone
# vectorstore = Pinecone.from_documents(documents=split_docs, embedding=embeddings, index_name="my-index")

# Qdrant (requires qdrant-client)
# from langchain_community.vectorstores import Qdrant
# vectorstore = Qdrant.from_documents(documents=split_docs, embedding=embeddings, location=":memory:")

The interface remains the same regardless of the backend, making it easy to switch between vector stores.

RETRIEVERS: FLEXIBLE DOCUMENT RETRIEVAL

Vector stores provide similarity search, but retrievers add additional functionality like filtering and hybrid search.

# Convert vector store to retriever
retriever = vectorstore.as_retriever(
    search_type="similarity",
    search_kwargs={"k": 2}
)

# Use the retriever
query = "How does memory work in LangChain?"
docs = retriever.invoke(query)

print(f"Query: {query}\n")
for doc in docs:
    print(f"Content: {doc.page_content}\n")

The as_retriever method creates a retriever from a vector store. Retrievers have a standard interface that works with LangChain's RAG chains.

You can configure different search types:

# Maximum Marginal Relevance (MMR) - balances relevance and diversity
mmr_retriever = vectorstore.as_retriever(
    search_type="mmr",
    search_kwargs={"k": 3, "fetch_k": 10}
)

docs = mmr_retriever.invoke("Tell me about LangChain features")

print("MMR Results:")
for i, doc in enumerate(docs, 1):
    print(f"{i}. {doc.page_content[:80]}...")

MMR retrieves more candidates (fetch_k) and then selects a diverse subset (k) that balances relevance and diversity. This prevents returning multiple very similar chunks.

You can also create custom retrievers with filtering:

from langchain_core.retrievers import BaseRetriever
from langchain_core.documents import Document
from typing import List

class MetadataFilterRetriever(BaseRetriever):
    """Retriever that filters by metadata."""
    
    vectorstore: FAISS
    metadata_filter: dict
    k: int = 3
    
    def _get_relevant_documents(self, query: str) -> List[Document]:
        """Retrieve documents matching the query and metadata filter."""
        # Get more candidates
        candidates = self.vectorstore.similarity_search(query, k=self.k * 3)
        
        # Filter by metadata
        filtered = [
            doc for doc in candidates
            if all(doc.metadata.get(key) == value 
                  for key, value in self.metadata_filter.items())
        ]
        
        # Return top k
        return filtered[:self.k]

# Use the custom retriever
filtered_retriever = MetadataFilterRetriever(
    vectorstore=vectorstore,
    metadata_filter={"topic": "concepts"},
    k=2
)

docs = filtered_retriever.invoke("Explain chains")
print("Filtered results (only 'concepts' topic):")
for doc in docs:
    print(f"Topic: {doc.metadata['topic']}")
    print(f"Content: {doc.page_content[:100]}...\n")

This custom retriever first retrieves candidates, then filters by metadata, and finally returns the top results. This pattern is useful when you want to restrict retrieval to specific document types or sources.

BUILDING RAG WITH LANGCHAIN

Now we can build a complete RAG system. LangChain provides high-level chains that handle the entire RAG workflow.

from langchain.chains import RetrievalQA
from langchain_openai import ChatOpenAI

# Create components
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
retriever = vectorstore.as_retriever(search_kwargs={"k": 2})

# Create RAG chain
qa_chain = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="stuff",
    retriever=retriever,
    return_source_documents=True
)

# Ask questions
query = "What are the different types of memory in LangChain?"
result = qa_chain.invoke({"query": query})

print(f"Question: {query}\n")
print(f"Answer: {result['result']}\n")
print("Source documents:")
for i, doc in enumerate(result['source_documents'], 1):
    print(f"{i}. {doc.page_content[:100]}...")

The RetrievalQA chain handles everything: retrieving relevant documents, formatting them into a prompt, calling the LLM, and returning the answer. The return_source_documents flag includes the retrieved chunks in the result.

The chain_type parameter controls how documents are combined. The "stuff" type puts all documents into a single prompt. Let's explore other types:

# Map-reduce: Process each document separately, then combine
mapreduce_chain = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="map_reduce",
    retriever=retriever
)

# Refine: Iteratively refine the answer with each document
refine_chain = RetrievalQA.from_chain_type(
    llm=llm,
    chain_type="refine",
    retriever=retriever
)

The "map_reduce" type processes each document independently and then combines the results. This is useful for long documents that don't fit in a single prompt. The "refine" type starts with an initial answer and refines it with each additional document.

For more control, use the modern LCEL (LangChain Expression Language) approach:

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

# Define the prompt template
template = """Answer the question based only on the following context:

{context}

Question: {question}

Answer:"""

prompt = ChatPromptTemplate.from_template(template)

# Create the RAG chain
def format_docs(docs):
    return "\n\n".join(doc.page_content for doc in docs)

rag_chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | prompt
    | llm
    | StrOutputParser()
)

# Use the chain
answer = rag_chain.invoke("How do agents work in LangChain?")
print(answer)

This LCEL chain is more explicit about what happens at each step. The retriever gets relevant documents, format_docs combines them into a string, the prompt template creates the final prompt, the LLM generates an answer, and the output parser extracts the text.

Let's add conversation memory to our RAG system:

from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory

# Create memory
memory = ConversationBufferMemory(
    memory_key="chat_history",
    return_messages=True,
    output_key="answer"
)

# Create conversational RAG chain
conversational_chain = ConversationalRetrievalChain.from_llm(
    llm=llm,
    retriever=retriever,
    memory=memory,
    return_source_documents=True
)

# Have a conversation
result1 = conversational_chain.invoke({"question": "What is LangChain?"})
print(f"Q: What is LangChain?")
print(f"A: {result1['answer']}\n")

result2 = conversational_chain.invoke({"question": "When was it created?"})
print(f"Q: When was it created?")
print(f"A: {result2['answer']}\n")

result3 = conversational_chain.invoke({"question": "Who created it?"})
print(f"Q: Who created it?")
print(f"A: {result3['answer']}")

The ConversationalRetrievalChain maintains conversation history and uses it to reformulate queries. When you ask "When was it created?", the chain understands that "it" refers to LangChain from the previous question.

AGENTS: AUTONOMOUS REASONING AND TOOL USE

Agents represent LangChain's most advanced capability: giving LLMs the ability to use tools and make decisions autonomously.

from langchain.agents import AgentExecutor, create_react_agent
from langchain.tools import Tool
from langchain_core.prompts import PromptTemplate

# Define some simple tools
def calculator(expression: str) -> str:
    """Evaluate a mathematical expression."""
    try:
        result = eval(expression)
        return f"The result is {result}"
    except Exception as e:
        return f"Error: {str(e)}"

def word_counter(text: str) -> str:
    """Count the number of words in text."""
    count = len(text.split())
    return f"The text contains {count} words"

# Create tool objects
tools = [
    Tool(
        name="Calculator",
        func=calculator,
        description="Useful for mathematical calculations. Input should be a valid Python expression."
    ),
    Tool(
        name="WordCounter",
        func=word_counter,
        description="Counts the number of words in a text. Input should be the text to count."
    )
]

# Create the agent prompt
agent_prompt = PromptTemplate.from_template(
    """Answer the following questions as best you can. You have access to the following tools:

{tools}

Use the following format:

Question: the input question you must answer Thought: you should always think about what to do Action: the action to take, should be one of [{tool_names}] Action Input: the input to the action Observation: the result of the action ... (this Thought/Action/Action Input/Observation can repeat N times) Thought: I now know the final answer Final Answer: the final answer to the original input question

Begin!

Question: {input} Thought: {agent_scratchpad}""" )

# Create the agent
agent = create_react_agent(llm, tools, agent_prompt)

# Create agent executor
agent_executor = AgentExecutor(
    agent=agent,
    tools=tools,
    verbose=True,
    max_iterations=5
)

# Use the agent
result = agent_executor.invoke({
    "input": "What is 25 * 17, and how many words are in this sentence?"
})

print(f"\nFinal Answer: {result['output']}")

The agent follows a reasoning loop called ReAct (Reasoning and Acting). It thinks about what to do, chooses a tool, executes it, observes the result, and repeats until it has the final answer.

The verbose flag shows the agent's thought process. You'll see it reason about which tools to use and how to combine their results.

Let's create a more practical agent with a search tool:

from langchain_community.tools import DuckDuckGoSearchRun

# Create a search tool
search = DuckDuckGoSearchRun()

# Wrap it as a LangChain tool
search_tool = Tool(
    name="Search",
    func=search.run,
    description="Useful for finding current information about topics. Input should be a search query."
)

# Create agent with search capability
search_tools = [search_tool, tools[0]]  # Search and calculator

search_agent = create_react_agent(llm, search_tools, agent_prompt)
search_executor = AgentExecutor(
    agent=search_agent,
    tools=search_tools,
    verbose=True
)

# Ask a question requiring search
result = search_executor.invoke({
    "input": "What is the current population of Tokyo, and what is that number divided by 1000?"
})

print(f"\nAnswer: {result['output']}")

The agent searches for Tokyo's population, then uses the calculator to divide it. This demonstrates how agents can chain multiple tools together to accomplish complex tasks.

You can create custom tools for any functionality:

from langchain.tools import BaseTool
from typing import Optional

class DocumentSearchTool(BaseTool):
    """Tool for searching the document knowledge base."""
    
    name = "DocumentSearch"
    description = "Search the LangChain documentation. Input should be a question about LangChain."
    retriever = None
    
    def __init__(self, retriever):
        super().__init__()
        self.retriever = retriever
    
    def _run(self, query: str) -> str:
        """Search the documents."""
        docs = self.retriever.invoke(query)
        if not docs:
            return "No relevant information found."
        
        result = "Found the following information:\n\n"
        for i, doc in enumerate(docs, 1):
            result += f"{i}. {doc.page_content}\n\n"
        return result
    
    async def _arun(self, query: str) -> str:
        """Async version."""
        raise NotImplementedError("Async not implemented")

# Create the tool
doc_search_tool = DocumentSearchTool(retriever=retriever)

# Create agent with document search
knowledge_tools = [doc_search_tool, tools[0]]
knowledge_agent = create_react_agent(llm, knowledge_tools, agent_prompt)
knowledge_executor = AgentExecutor(
    agent=knowledge_agent,
    tools=knowledge_tools,
    verbose=True
)

# Use it
result = knowledge_executor.invoke({
    "input": "What types of memory does LangChain support, and how many are there?"
})

print(f"\nAnswer: {result['output']}")

This agent can search your document knowledge base and perform calculations, combining retrieval with reasoning.

ADVANCED AGENT PATTERNS

LangChain supports more sophisticated agent architectures. The OpenAI Functions agent uses function calling for more reliable tool use:

from langchain.agents import create_openai_functions_agent

# Create a more structured prompt
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder

functions_prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant with access to tools."),
    ("human", "{input}"),
    MessagesPlaceholder(variable_name="agent_scratchpad")
])

# Create OpenAI functions agent
functions_agent = create_openai_functions_agent(llm, tools, functions_prompt)
functions_executor = AgentExecutor(
    agent=functions_agent,
    tools=tools,
    verbose=True
)

# Use it
result = functions_executor.invoke({
    "input": "Calculate 15 squared, then count the words in 'LangChain makes building AI applications easier'"
})

print(f"\nResult: {result['output']}")

The OpenAI Functions agent is more reliable because it uses the model's built-in function calling capability rather than parsing text output.

You can also create agents with memory:

from langchain.agents import AgentExecutor, create_react_agent
from langchain.memory import ConversationBufferMemory

# Create memory for the agent
agent_memory = ConversationBufferMemory(
    memory_key="chat_history",
    return_messages=True
)

# Create agent with memory
memory_agent_prompt = PromptTemplate.from_template(
    """Answer questions using available tools. You have access to:

{tools}

Previous conversation: {chat_history}

Current question: {input}

{agent_scratchpad}""" )

# Note: Integrating memory with agents requires careful prompt engineering
# This is a simplified example

This allows agents to remember previous interactions and use that context in decision-making.

PRODUCTION CONSIDERATIONS

When deploying LangChain applications to production, several considerations become important.

First is error handling. LangChain operations can fail for many reasons: API rate limits, network issues, invalid inputs, or model errors.

from langchain.callbacks import get_openai_callback
import time

def robust_chain_invoke(chain, input_data, max_retries=3):
    """Invoke a chain with retry logic and error handling."""
    for attempt in range(max_retries):
        try:
            with get_openai_callback() as cb:
                result = chain.invoke(input_data)
                
                # Log token usage
                print(f"Tokens used: {cb.total_tokens}")
                print(f"Cost: ${cb.total_cost:.4f}")
                
                return result
                
        except Exception as e:
            print(f"Attempt {attempt + 1} failed: {str(e)}")
            
            if attempt < max_retries - 1:
                # Exponential backoff
                wait_time = 2 ** attempt
                print(f"Retrying in {wait_time} seconds...")
                time.sleep(wait_time)
            else:
                print("Max retries reached")
                raise

# Use the robust invoke
try:
    result = robust_chain_invoke(rag_chain, "What is LangChain?")
    print(f"Result: {result}")
except Exception as e:
    print(f"Failed after retries: {e}")

The get_openai_callback context manager tracks token usage and costs, which is crucial for monitoring production applications.

Second is caching. Repeated queries should use cached results to save time and money:

from langchain.cache import InMemoryCache
from langchain.globals import set_llm_cache

# Enable caching
set_llm_cache(InMemoryCache())

# Now LLM calls are cached
llm = ChatOpenAI(model="gpt-3.5-turbo")

# First call - hits the API
start = time.time()
result1 = llm.invoke("What is the capital of France?")
time1 = time.time() - start

# Second call - uses cache
start = time.time()
result2 = llm.invoke("What is the capital of France?")
time2 = time.time() - start

print(f"First call: {time1:.2f}s")
print(f"Second call: {time2:.2f}s (cached)")

For persistent caching across sessions, use SQLite or Redis:

from langchain.cache import SQLiteCache

# Use SQLite cache
set_llm_cache(SQLiteCache(database_path=".langchain.db"))

Third is monitoring and logging. Production applications need observability:

from langchain.callbacks import StdOutCallbackHandler
from langchain.callbacks.base import BaseCallbackHandler

class CustomCallbackHandler(BaseCallbackHandler):
    """Custom callback for logging."""
    
    def on_llm_start(self, serialized, prompts, **kwargs):
        """Log when LLM starts."""
        print(f"[LLM START] Prompts: {len(prompts)}")
    
    def on_llm_end(self, response, **kwargs):
        """Log when LLM ends."""
        print(f"[LLM END] Tokens: {response.llm_output.get('token_usage', {})}")
    
    def on_chain_start(self, serialized, inputs, **kwargs):
        """Log when chain starts."""
        print(f"[CHAIN START] {serialized.get('name', 'Unknown')}")
    
    def on_chain_end(self, outputs, **kwargs):
        """Log when chain ends."""
        print(f"[CHAIN END]")

# Use the callback
callbacks = [CustomCallbackHandler()]

chain = prompt | llm | StrOutputParser()
result = chain.invoke(
    {"topic": "machine learning"},
    config={"callbacks": callbacks}
)

Callbacks provide hooks into LangChain's execution flow, allowing you to log, monitor, and debug your applications.

Fourth is streaming for better user experience:

from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler

# Create streaming LLM
streaming_llm = ChatOpenAI(
    model="gpt-3.5-turbo",
    streaming=True,
    callbacks=[StreamingStdOutCallbackHandler()]
)

# Use in a chain
streaming_chain = prompt | streaming_llm | StrOutputParser()

print("Streaming response:")
result = streaming_chain.invoke({"topic": "quantum computing"})

The response streams token by token, providing immediate feedback to users.

For RAG systems, you can stream both retrieval and generation:

from langchain.callbacks.manager import CallbackManager

class StreamingRAGHandler(BaseCallbackHandler):
    """Handler for streaming RAG responses."""
    
    def on_retriever_end(self, documents, **kwargs):
        """Called when retrieval completes."""
        print(f"\n[Retrieved {len(documents)} documents]\n")
    
    def on_llm_new_token(self, token: str, **kwargs):
        """Called for each new token."""
        print(token, end="", flush=True)

# Create streaming RAG chain
streaming_rag = ConversationalRetrievalChain.from_llm(
    llm=ChatOpenAI(
        model="gpt-3.5-turbo",
        streaming=True,
        callbacks=[StreamingRAGHandler()]
    ),
    retriever=retriever,
    memory=ConversationBufferMemory(
        memory_key="chat_history",
        return_messages=True,
        output_key="answer"
    )
)

result = streaming_rag.invoke({"question": "What are agents in LangChain?"})

LANGCHAIN VS DIRECT IMPLEMENTATION

When should you use LangChain versus implementing with lower-level libraries like HuggingFace?

LangChain excels when you need rapid prototyping. Building a RAG system with LangChain takes minutes instead of hours. The abstractions handle common patterns, letting you focus on application logic.

LangChain is ideal for standard workflows. If your use case fits LangChain's patterns (chatbots, RAG, agents), you benefit from battle-tested implementations and community support.

LangChain provides ecosystem integration. It connects to hundreds of services: vector databases, LLM providers, document loaders, and tools. This integration is valuable for complex applications.

However, direct implementation offers more control. You can optimize every detail for your specific use case. You avoid the abstraction overhead and potential bugs in the framework.

Direct implementation is better for learning. Understanding how RAG works at a low level makes you a better AI engineer. LangChain can be a black box that hides important details.

Direct implementation may be necessary for custom requirements. If your use case doesn't fit LangChain's patterns, fighting the framework can be harder than building from scratch.

The best approach often combines both. Use LangChain for rapid prototyping and standard components. Drop down to lower-level libraries for custom or performance-critical parts.

Here's an example mixing LangChain with custom code:

from langchain_community.vectorstores import FAISS
from langchain_community.embeddings import HuggingFaceEmbeddings
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

# Use LangChain for retrieval
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
vectorstore = FAISS.from_documents(split_docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})

# Use HuggingFace directly for generation (more control)
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")

def custom_rag_answer(question):
    """Custom RAG implementation mixing LangChain and HuggingFace."""
    # Use LangChain retriever
    docs = retriever.invoke(question)
    context = "\n\n".join(doc.page_content for doc in docs)
    
    # Custom prompt formatting
    prompt = f"Context:\n{context}\n\nQuestion: {question}\n\nAnswer:"
    
    # Use HuggingFace for generation with custom parameters
    inputs = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=512)
    
    with torch.no_grad():
        outputs = model.generate(
            inputs["input_ids"],
            max_length=len(inputs["input_ids"][0]) + 100,
            temperature=0.7,
            top_p=0.9,
            do_sample=True,
            pad_token_id=tokenizer.eos_token_id
        )
    
    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
    answer = response[len(prompt):].strip()
    
    return {
        "answer": answer,
        "context": context,
        "sources": [doc.metadata for doc in docs]
    }

# Use the hybrid approach
result = custom_rag_answer("What is LangChain?")
print(f"Answer: {result['answer']}")
print(f"\nSources: {result['sources']}")

This approach uses LangChain's retriever (well-tested, handles multiple vector stores) but custom generation logic (full control over parameters and prompt formatting).

ADVANCED LANGCHAIN PATTERNS

Let's explore some advanced patterns that showcase LangChain's power.

First is the multi-query retriever, which generates multiple search queries for better recall:

from langchain.retrievers.multi_query import MultiQueryRetriever

# Create multi-query retriever
multi_retriever = MultiQueryRetriever.from_llm(
    retriever=vectorstore.as_retriever(),
    llm=llm
)

# It generates multiple queries and combines results
docs = multi_retriever.invoke("How does LangChain handle conversations?")

print(f"Retrieved {len(docs)} documents")
for doc in docs:
    print(f"- {doc.page_content[:80]}...")

The multi-query retriever uses the LLM to generate alternative phrasings of the query, retrieves documents for each, and combines the results. This improves recall for ambiguous questions.

Second is the contextual compression retriever, which filters retrieved documents:

from langchain.retrievers import ContextualCompressionRetriever
from langchain.retrievers.document_compressors import LLMChainExtractor

# Create compressor
compressor = LLMChainExtractor.from_llm(llm)

# Create compression retriever
compression_retriever = ContextualCompressionRetriever(
    base_compressor=compressor,
    base_retriever=retriever
)

# Use it
compressed_docs = compression_retriever.invoke("What are chains?")

print("Compressed documents:")
for doc in compressed_docs:
    print(f"- {doc.page_content}")

The compressor uses the LLM to extract only the relevant parts of each document, reducing noise and improving answer quality.

Third is the parent document retriever, which retrieves small chunks but provides larger context:

from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore
from langchain.text_splitter import RecursiveCharacterTextSplitter

# Create storage for parent documents
store = InMemoryStore()

# Create splitters for parent and child chunks
parent_splitter = RecursiveCharacterTextSplitter(chunk_size=400)
child_splitter = RecursiveCharacterTextSplitter(chunk_size=100)

# Create parent document retriever
parent_retriever = ParentDocumentRetriever(
    vectorstore=vectorstore,
    docstore=store,
    child_splitter=child_splitter,
    parent_splitter=parent_splitter
)

# Add documents
parent_retriever.add_documents(documents)

# Retrieve - returns parent documents even though search uses child chunks
docs = parent_retriever.invoke("Explain memory systems")

print("Retrieved parent documents:")
for doc in docs:
    print(f"Length: {len(doc.page_content)} chars")
    print(f"Content: {doc.page_content[:100]}...\n")

This retriever searches using small chunks (for precision) but returns larger parent documents (for context). It's useful when you need both precise retrieval and sufficient context.

Fourth is the ensemble retriever, which combines multiple retrieval methods:

from langchain.retrievers import EnsembleRetriever
from langchain_community.retrievers import BM25Retriever

# Create BM25 retriever (keyword-based)
bm25_retriever = BM25Retriever.from_documents(split_docs)
bm25_retriever.k = 2

# Create ensemble combining semantic and keyword search
ensemble_retriever = EnsembleRetriever(
    retrievers=[retriever, bm25_retriever],
    weights=[0.5, 0.5]
)

# Use it
docs = ensemble_retriever.invoke("LangChain framework features")

print("Ensemble retrieval results:")
for doc in docs:
    print(f"- {doc.page_content[:80]}...")

The ensemble retriever combines semantic search (vector similarity) with keyword search (BM25), providing better results than either method alone.

BUILDING A COMPLETE APPLICATION

Let's bring everything together into a complete LangChain application: a conversational RAG chatbot with tool use.

from langchain_openai import ChatOpenAI
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain.agents import AgentExecutor, create_openai_functions_agent
from langchain.tools import Tool
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.documents import Document

class AdvancedChatbot:
    """A complete LangChain chatbot with RAG and tools."""
    
    def __init__(self, documents):
        """Initialize the chatbot."""
        print("Initializing chatbot...")
        
        # Create LLM
        self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
        
        # Process documents
        print("Processing documents...")
        text_splitter = RecursiveCharacterTextSplitter(
            chunk_size=200,
            chunk_overlap=50
        )
        self.chunks = text_splitter.split_documents(documents)
        
        # Create vector store
        print("Creating vector store...")
        embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
        self.vectorstore = FAISS.from_documents(self.chunks, embeddings)
        self.retriever = self.vectorstore.as_retriever(search_kwargs={"k": 3})
        
        # Create memory
        self.memory = ConversationBufferMemory(
            memory_key="chat_history",
            return_messages=True,
            output_key="answer"
        )
        
        # Create RAG chain
        self.rag_chain = ConversationalRetrievalChain.from_llm(
            llm=self.llm,
            retriever=self.retriever,
            memory=self.memory,
            return_source_documents=True
        )
        
        # Create tools
        self.tools = self._create_tools()
        
        # Create agent
        self.agent_executor = self._create_agent()
        
        print("Chatbot ready!")
    
    def _create_tools(self):
        """Create tools for the agent."""
        def search_docs(query: str) -> str:
            """Search the knowledge base."""
            docs = self.retriever.invoke(query)
            if not docs:
                return "No information found."
            return "\n\n".join(doc.page_content for doc in docs)
        
        def calculate(expression: str) -> str:
            """Evaluate a mathematical expression."""
            try:
                result = eval(expression)
                return str(result)
            except:
                return "Invalid expression"
        
        return [
            Tool(
                name="DocumentSearch",
                func=search_docs,
                description="Search the LangChain knowledge base. Use this for questions about LangChain."
            ),
            Tool(
                name="Calculator",
                func=calculate,
                description="Perform calculations. Input should be a mathematical expression."
            )
        ]
    
    def _create_agent(self):
        """Create the agent."""
        prompt = ChatPromptTemplate.from_messages([
            ("system", "You are a helpful assistant with access to tools."),
            ("human", "{input}"),
            MessagesPlaceholder(variable_name="agent_scratchpad")
        ])
        
        agent = create_openai_functions_agent(self.llm, self.tools, prompt)
        return AgentExecutor(agent=agent, tools=self.tools, verbose=True)
    
    def chat(self, message: str, use_agent: bool = False):
        """Chat with the bot."""
        if use_agent:
            # Use agent for complex queries
            result = self.agent_executor.invoke({"input": message})
            return result["output"]
        else:
            # Use RAG chain for document questions
            result = self.rag_chain.invoke({"question": message})
            return {
                "answer": result["answer"],
                "sources": [doc.page_content[:100] for doc in result["source_documents"]]
            }
    
    def reset(self):
        """Reset conversation memory."""
        self.memory.clear()

# Create and use the chatbot
chatbot = AdvancedChatbot(documents)

# Ask RAG questions
print("\n=== RAG Mode ===")
response = chatbot.chat("What is LangChain?")
print(f"Answer: {response['answer']}")
print(f"Sources: {response['sources']}\n")

response = chatbot.chat("When was it created?")
print(f"Answer: {response['answer']}\n")

# Use agent for complex tasks
print("\n=== Agent Mode ===")
response = chatbot.chat(
    "Search for information about chains, then calculate 10 * 5",
    use_agent=True
)
print(f"Answer: {response}")

This complete application demonstrates LangChain's power: conversational RAG, memory management, tool use, and agent reasoning, all in a clean, maintainable structure.

WHAT YOU'VE MASTERED

Congratulations! You've journeyed from LangChain basics to building sophisticated AI applications. Let's review what you now understand.

You learned LangChain's core philosophy: composability, abstraction, standardization, and extensibility. You understand how these principles guide the framework's design.

You mastered the fundamental components. You know how to use language models with consistent interfaces. You understand prompt templates for dynamic prompt engineering. You can create chains that compose operations into workflows.

You explored memory systems for maintaining conversation context. You understand the tradeoffs between different memory types and when to use each.

You learned document processing: loading documents from various sources, splitting them intelligently, and creating searchable vector stores. You understand retrievers and their different search strategies.

You built complete RAG systems using LangChain's high-level chains. You understand how retrieval, context injection, and generation work together.

You discovered agents and tool use. You know how agents reason about which tools to use and how to create custom tools for specific tasks.

You explored advanced patterns: multi-query retrieval, contextual compression, parent document retrieval, and ensemble retrieval. You understand when each pattern is appropriate.

You learned production considerations: error handling, caching, monitoring, and streaming. You know how to build robust, observable applications.

Most importantly, you understand when to use LangChain and when to use lower-level libraries. You can make informed architectural decisions.

YOUR NEXT STEPS

Your LangChain journey continues beyond this tutorial. Here are directions to explore.

Build real applications. Create a chatbot for your company's documentation. Build a research assistant that searches papers and synthesizes findings. Create a code assistant that understands your codebase. Real projects teach lessons tutorials can't.

Explore the LangChain ecosystem. Try LangSmith for debugging and monitoring. Experiment with LangServe for deploying chains as APIs. Use LangGraph for building complex, stateful agents.

Study advanced agent patterns. Learn about plan-and-execute agents, multi-agent systems, and hierarchical agents. These patterns enable more sophisticated autonomous behavior.

Experiment with different LLM providers. Try Anthropic's Claude, Google's PaLM, or open-source models through Ollama. Compare their strengths and weaknesses.

Dive into prompt engineering. Learn techniques like chain-of-thought, tree-of-thought, and self-consistency. Master few-shot learning and instruction tuning.

Contribute to the community. LangChain is open source. Report bugs, suggest features, or contribute code. The community is active and welcoming.

Stay current with research. The field evolves rapidly. Follow papers on arXiv, read the LangChain blog, and join discussions on Discord and forums.

Optimize for production. Learn about model quantization, caching strategies, and deployment architectures. Understand the economics of running LLM applications at scale.

RESOURCES FOR CONTINUED LEARNING

The official LangChain documentation is comprehensive and regularly updated. Start with the conceptual guides to deepen your understanding.

The LangChain cookbook provides practical recipes for common tasks. It's an excellent resource for learning patterns and best practices.

LangChain's YouTube channel has tutorials and talks from the team and community. Visual learning complements written documentation.

The LangChain Discord server is active and helpful. Ask questions, share projects, and learn from others building with LangChain.

Harrison Chase's blog posts and talks provide insights into LangChain's design philosophy and future direction.

Papers like ReAct, Tree of Thoughts, and Reflexion explain the reasoning patterns that agents use. Understanding these papers helps you build better agents.

The Awesome LangChain repository curates tools, tutorials, and projects. It's a great way to discover what's possible.

FINAL REFLECTIONS

LangChain represents a shift in how we build AI applications. Instead of writing procedural code that calls APIs, we compose declarative chains that express intent. Instead of managing state manually, we use memory systems. Instead of hard-coding logic, we create agents that reason.

This abstraction has costs and benefits. The costs include learning curve, abstraction overhead, and reduced control. The benefits include rapid development, battle-tested patterns, and ecosystem integration.

The key is knowing when to use which approach. For standard workflows like chatbots and RAG, LangChain accelerates development dramatically. For custom requirements or performance-critical code, lower-level libraries provide necessary control. The best applications often mix both.

You now have the knowledge to build sophisticated AI applications with LangChain. You understand the abstractions, the patterns, and the tradeoffs. You can create chatbots, RAG systems, and autonomous agents.

But more than specific techniques, you've gained a mental model of how to architect AI applications. You understand composition, abstraction, and orchestration. These concepts transcend any particular framework.

The AI field is advancing rapidly. New models, new techniques, and new frameworks emerge constantly. But the fundamentals you've learned here will serve you well. Understanding retrieval, generation, reasoning, and tool use prepares you for whatever comes next.

Keep building. Keep learning. Keep experimenting. The future of AI applications is being written right now, and you're equipped to be part of it.

Welcome to the world of LangChain development. Now go build something extraordinary.