THE ORCHESTRATION LAYER FOR AI APPLICATIONS
Imagine you're building a house. You could craft every nail, cut every board, and mix every batch of concrete yourself. Or you could use pre-made materials and focus on the architecture. LangChain is the latter approach for AI applications.
When you built the HuggingFace chatbot in the previous tutorial, you manually connected components: loading models, creating embeddings, managing vector stores, formatting prompts, and orchestrating retrieval. It worked, but it required understanding and implementing every detail. LangChain provides a higher-level abstraction layer that handles these patterns for you.
Founded by Harrison Chase in late 2022, LangChain exploded in popularity because it solved a real problem: building LLM applications involves repetitive patterns. Every RAG system needs document loading, text splitting, embedding, vector storage, and retrieval. Every chatbot needs memory management. Every agent needs tool integration. LangChain provides battle-tested implementations of these patterns.
But LangChain is more than a convenience library. It introduces powerful abstractions that change how you think about AI applications. Instead of writing procedural code that calls APIs, you compose declarative chains that express what you want to happen. Instead of managing state manually, you use memory systems. Instead of hard-coding logic, you create agents that reason about which tools to use.
By the end of this tutorial, you'll understand LangChain's core concepts and abstractions. You'll build chatbots with memory, implement RAG systems with just a few lines of code, and create agents that can use tools autonomously. You'll understand when to use LangChain and when lower-level libraries like HuggingFace are more appropriate.
Let's dive into the world of LangChain and discover how it transforms AI application development.
UNDERSTANDING LANGCHAIN'S PHILOSOPHY
Before writing code, let's understand LangChain's design philosophy. This will help you think in "LangChain terms" and use the library effectively.
LangChain is built around several core principles. First is composability. Complex applications are built by composing simple components. A RAG system composes a retriever with a language model. An agent composes tools with a reasoning engine. This compositional approach makes systems easier to understand, test, and modify.
Second is abstraction. LangChain provides abstract interfaces for common components. Whether you use OpenAI, Anthropic, or a local model, the interface remains the same. Whether you store vectors in FAISS, Pinecone, or Chroma, the retriever interface is consistent. This abstraction lets you swap components without rewriting your application.
Third is standardization. LangChain establishes patterns for common tasks. There's a standard way to handle chat history, a standard way to structure prompts, a standard way to implement retrieval. These patterns emerge from real-world usage and represent best practices.
Fourth is extensibility. While LangChain provides many built-in components, you can easily create custom ones. Need a special document loader? Implement the base class. Need a custom tool for your agent? Define its interface. The framework is designed for extension.
Understanding these principles helps you use LangChain effectively. You're not just calling functions; you're composing components into systems.
SETTING UP YOUR LANGCHAIN ENVIRONMENT
Let's prepare our development environment. LangChain has a modular structure, so we'll install the components we need.
pip install langchain langchain-community langchain-openai
pip install faiss-cpu sentence-transformers
pip install python-dotenv
The langchain package contains the core abstractions and base classes. The langchain-community package includes integrations with various services and tools. The langchain-openai package provides OpenAI integrations. We'll also install FAISS for vector storage and sentence-transformers for embeddings.
LangChain works with many LLM providers. For this tutorial, we'll use OpenAI's API because it's reliable and well-documented. You'll need an API key from OpenAI. Create a file named dot-env in your project directory:
OPENAI_API_KEY=your-api-key-here
Now let's verify the installation:
import langchain
from langchain_openai import ChatOpenAI
from dotenv import load_dotenv
import os
# Load environment variables
load_dotenv()
# Verify API key is loaded
api_key = os.getenv("OPENAI_API_KEY")
if api_key:
print("API key loaded successfully")
print(f"LangChain version: {langchain.__version__}")
else:
print("Warning: OPENAI_API_KEY not found in environment")
If you prefer to use local models instead of OpenAI, you can use Ollama or HuggingFace models. We'll show alternatives throughout the tutorial.
YOUR FIRST LANGCHAIN INTERACTION
Let's start with the simplest possible LangChain program: asking a question to an LLM.
from langchain_openai import ChatOpenAI
from dotenv import load_dotenv
# Load environment variables
load_dotenv()
# Create a language model instance
llm = ChatOpenAI(
model="gpt-3.5-turbo",
temperature=0.7
)
# Ask a question
response = llm.invoke("What are the three laws of robotics?")
print(response.content)
This code creates a ChatOpenAI instance, which is LangChain's wrapper around OpenAI's chat models. The invoke method sends a message and returns the response. The temperature parameter controls randomness, just like in the HuggingFace tutorial.
The response object contains more than just text. Let's explore it:
response = llm.invoke("Explain quantum computing in one sentence")
print(f"Content: {response.content}")
print(f"Response metadata: {response.response_metadata}")
print(f"Type: {type(response)}")
The response is an AIMessage object containing the content, metadata about the API call (like token usage), and other information. This structured response makes it easy to extract what you need.
Now let's see LangChain's real power: handling conversations with multiple messages.
from langchain_core.messages import HumanMessage, SystemMessage, AIMessage
# Create a conversation
messages = [
SystemMessage(content="You are a helpful physics tutor."),
HumanMessage(content="What is quantum entanglement?"),
]
response = llm.invoke(messages)
print(f"Assistant: {response.content}\n")
# Continue the conversation
messages.append(response)
messages.append(HumanMessage(content="Can you give me a simple analogy?"))
response = llm.invoke(messages)
print(f"Assistant: {response.content}")
LangChain uses message objects to represent different roles in a conversation. SystemMessage sets the AI's behavior and context. HumanMessage represents user input. AIMessage represents the AI's responses. This structure mirrors how chat models actually work.
The conversation maintains context because we pass the entire message history with each call. The model sees the previous exchange and can provide coherent follow-up responses.
PROMPT TEMPLATES: STRUCTURED PROMPT ENGINEERING
Hard-coding prompts works for simple cases, but real applications need dynamic prompts that incorporate variables. LangChain's prompt templates solve this elegantly.
from langchain_core.prompts import ChatPromptTemplate
# Create a prompt template
template = ChatPromptTemplate.from_messages([
("system", "You are a helpful assistant that translates {input_language} to {output_language}."),
("human", "{text}")
])
# Format the prompt with variables
messages = template.format_messages(
input_language="English",
output_language="French",
text="Hello, how are you?"
)
# Use with the LLM
response = llm.invoke(messages)
print(response.content)
The template uses curly braces for variables. When you call format_messages, it substitutes the variables with actual values. This separation of template and data makes prompts reusable and testable.
You can create more complex templates with multiple variables and conditional logic:
from langchain_core.prompts import PromptTemplate
# Template for code explanation
code_template = PromptTemplate(
input_variables=["language", "code", "detail_level"],
template="""You are an expert {language} programmer.
Explain the following code at a {detail_level} level of detail.
Code: {code}
Explanation:""" )
# Use the template
prompt = code_template.format(
language="Python",
detail_level="beginner-friendly",
code="def fibonacci(n):\n return n if n <= 1 else fibonacci(n-1) + fibonacci(n-2)"
)
response = llm.invoke(prompt)
print(response.content)
Templates support different formats for different use cases. ChatPromptTemplate is for chat models, while PromptTemplate is for completion models. There are also specialized templates for few-shot learning and other patterns.
Here's a practical example showing few-shot prompting:
from langchain_core.prompts import FewShotPromptTemplate, PromptTemplate
# Examples for few-shot learning
examples = [
{
"input": "The movie was fantastic!",
"output": "Positive"
},
{
"input": "I hated every minute of it.",
"output": "Negative"
},
{
"input": "It was okay, nothing special.",
"output": "Neutral"
}
]
# Template for each example
example_template = PromptTemplate(
input_variables=["input", "output"],
template="Input: {input}\nSentiment: {output}"
)
# Few-shot prompt template
few_shot_template = FewShotPromptTemplate(
examples=examples,
example_prompt=example_template,
prefix="Classify the sentiment of the following text.",
suffix="Input: {input}\nSentiment:",
input_variables=["input"]
)
# Use the template
prompt = few_shot_template.format(input="This product exceeded my expectations!")
response = llm.invoke(prompt)
print(response.content)
Few-shot prompting provides examples to guide the model's behavior. This is especially useful for tasks where you want consistent output formatting or specific classification categories.
CHAINS: COMPOSING OPERATIONS
Now we reach one of LangChain's most powerful concepts: chains. A chain is a sequence of operations that process data. The output of one step becomes the input to the next.
The simplest chain connects a prompt template to an LLM:
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
# Create components
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
prompt = ChatPromptTemplate.from_messages([
("system", "You are a creative storyteller."),
("human", "Write a one-paragraph story about {topic}")
])
# Create a chain using the pipe operator
chain = prompt | llm
# Invoke the chain
response = chain.invoke({"topic": "a robot learning to paint"})
print(response.content)
The pipe operator (vertical bar) creates a chain. When you invoke the chain with a dictionary of variables, it flows through each component. The prompt template formats the messages, then the LLM generates a response.
You can extend chains with output parsers to structure the response:
from langchain_core.output_parsers import StrOutputParser
# Add an output parser to extract just the string content
chain = prompt | llm | StrOutputParser()
# Now the result is a string, not an AIMessage object
result = chain.invoke({"topic": "a time-traveling historian"})
print(f"Type: {type(result)}")
print(f"Story: {result}")
The StrOutputParser extracts the content string from the AIMessage. This makes the chain's output cleaner and easier to use in subsequent operations.
Let's build a more complex chain that performs multiple steps:
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain_core.output_parsers import StrOutputParser
# First chain: Generate a topic
topic_prompt = ChatPromptTemplate.from_template(
"Suggest an interesting topic for a {genre} story. Just give the topic, nothing else."
)
topic_chain = topic_prompt | llm | StrOutputParser()
# Second chain: Write the story
story_prompt = ChatPromptTemplate.from_template(
"Write a short {genre} story about: {topic}"
)
story_chain = story_prompt | llm | StrOutputParser()
# Combine the chains
def generate_story(genre):
"""Generate a story by first creating a topic, then writing about it."""
# Get a topic
topic = topic_chain.invoke({"genre": genre})
print(f"Generated topic: {topic}\n")
# Write the story
story = story_chain.invoke({"genre": genre, "topic": topic})
return story
# Use the combined workflow
result = generate_story("science fiction")
print(result)
This demonstrates sequential processing where the output of one chain feeds into another. The first chain generates a topic, and the second chain uses that topic to write a story.
LangChain provides specialized chain types for common patterns. The LLMChain is a simple prompt-to-LLM chain:
from langchain.chains import LLMChain
# Create an LLMChain
prompt = ChatPromptTemplate.from_template("What is the capital of {country}?")
chain = LLMChain(llm=llm, prompt=prompt)
# Use it
result = chain.invoke({"country": "Japan"})
print(result["text"])
While the modern approach uses the pipe operator, LLMChain is still useful for backward compatibility and certain use cases.
MEMORY: MAINTAINING CONVERSATION CONTEXT
Real chatbots need to remember previous exchanges. LangChain's memory systems handle this automatically.
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationChain
# Create a memory instance
memory = ConversationBufferMemory()
# Create a conversation chain with memory
conversation = ConversationChain(
llm=llm,
memory=memory,
verbose=True # Shows what's happening internally
)
# Have a conversation
print(conversation.predict(input="Hi, my name is Alice"))
print("\n" + "="*50 + "\n")
print(conversation.predict(input="What's 2+2?"))
print("\n" + "="*50 + "\n")
print(conversation.predict(input="What's my name?"))
The ConversationBufferMemory stores all messages in a buffer. When you ask "What's my name?", the model can answer because it has access to the entire conversation history.
The verbose flag shows you what's being sent to the LLM. You'll see that each call includes the full conversation history, allowing the model to maintain context.
Let's examine the memory directly:
# View the conversation history
print("\nConversation history:")
print(memory.load_memory_variables({}))
# Clear the memory
memory.clear()
print("\nMemory cleared")
The load_memory_variables method returns the stored conversation. This is useful for debugging or saving conversations.
LangChain offers different memory types for different needs. ConversationBufferWindowMemory keeps only the last N messages:
from langchain.memory import ConversationBufferWindowMemory
# Keep only the last 3 exchanges
windowed_memory = ConversationBufferWindowMemory(k=3)
conversation = ConversationChain(
llm=llm,
memory=windowed_memory
)
# Have a longer conversation
conversation.predict(input="My favorite color is blue")
conversation.predict(input="I work as a software engineer")
conversation.predict(input="I enjoy hiking on weekends")
conversation.predict(input="I have a cat named Whiskers")
# This will only remember the last 3 exchanges
response = conversation.predict(input="What's my favorite color?")
print(response)
The model might not remember your favorite color because that exchange has been pushed out of the window. This memory type is useful for long conversations where you want to limit token usage.
Another useful memory type is ConversationSummaryMemory, which summarizes old messages:
from langchain.memory import ConversationSummaryMemory
# Create summary memory
summary_memory = ConversationSummaryMemory(llm=llm)
conversation = ConversationChain(
llm=llm,
memory=summary_memory,
verbose=True
)
# Have a conversation
conversation.predict(input="I'm planning a trip to Japan next spring")
conversation.predict(input="I want to visit Tokyo, Kyoto, and Osaka")
conversation.predict(input="I'm particularly interested in traditional temples")
# Check the summary
print("\nMemory summary:")
print(summary_memory.load_memory_variables({}))
The summary memory uses the LLM to create a running summary of the conversation. This keeps token usage low while maintaining important context.
For more control, you can use ConversationBufferMemory with custom keys:
from langchain.memory import ConversationBufferMemory
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
# Create memory with custom key
memory = ConversationBufferMemory(
memory_key="chat_history",
return_messages=True
)
# Create a prompt that uses the memory
prompt = ChatPromptTemplate.from_messages([
("system", "You are a helpful assistant."),
MessagesPlaceholder(variable_name="chat_history"),
("human", "{input}")
])
# Create chain with memory
from langchain.chains import LLMChain
chain = LLMChain(
llm=llm,
prompt=prompt,
memory=memory
)
# Use the chain
response = chain.predict(input="Hi, I'm learning LangChain")
print(response)
response = chain.predict(input="What am I learning?")
print(response)
The MessagesPlaceholder reserves a spot in the prompt for the conversation history. This gives you fine-grained control over where history appears in your prompts.
DOCUMENT LOADERS: INGESTING INFORMATION
Now we're ready to build RAG systems with LangChain. The first step is loading documents. LangChain provides loaders for many file types.
from langchain_community.document_loaders import TextLoader
# Load a text file
loader = TextLoader("example.txt")
documents = loader.load()
print(f"Loaded {len(documents)} document(s)")
print(f"First document preview: {documents[0].page_content[:200]}")
print(f"Metadata: {documents[0].metadata}")
Each document has page_content (the text) and metadata (information about the source). The metadata typically includes the filename and other relevant details.
For this tutorial, let's create documents programmatically:
from langchain_core.documents import Document
# Create sample documents about LangChain
documents = [
Document(
page_content="""LangChain is a framework for developing applications powered by
language models. It was created by Harrison Chase and released in October 2022.
The framework provides abstractions for working with LLMs, including chains,
agents, and memory systems.""",
metadata={"source": "intro", "topic": "overview"}
),
Document(
page_content="""Chains in LangChain are sequences of operations that process data.
The simplest chain connects a prompt template to an LLM. More complex chains can
include multiple steps, conditional logic, and parallel execution. Chains are
composable, meaning you can combine simple chains into complex workflows.""",
metadata={"source": "chains", "topic": "concepts"}
),
Document(
page_content="""Memory systems in LangChain maintain conversation context.
ConversationBufferMemory stores all messages. ConversationBufferWindowMemory
keeps only recent messages. ConversationSummaryMemory creates summaries of old
messages. Each memory type offers different tradeoffs between context preservation
and token usage.""",
metadata={"source": "memory", "topic": "concepts"}
),
Document(
page_content="""Agents in LangChain can use tools to accomplish tasks. An agent
receives a task, reasons about which tools to use, executes the tools, and
synthesizes the results. Tools can be anything from calculators to search engines
to custom APIs. Agents enable autonomous behavior where the LLM decides what
actions to take.""",
metadata={"source": "agents", "topic": "advanced"}
)
]
print(f"Created {len(documents)} documents")
LangChain supports many document loaders. Here are some examples:
# PDF loader (requires pypdf)
# from langchain_community.document_loaders import PyPDFLoader
# loader = PyPDFLoader("document.pdf")
# documents = loader.load()
# Web page loader (requires beautifulsoup4)
# from langchain_community.document_loaders import WebBaseLoader
# loader = WebBaseLoader("https://example.com")
# documents = loader.load()
# CSV loader
# from langchain_community.document_loaders import CSVLoader
# loader = CSVLoader("data.csv")
# documents = loader.load()
# Directory loader (loads all files in a directory)
# from langchain_community.document_loaders import DirectoryLoader
# loader = DirectoryLoader("./documents", glob="**/*.txt")
# documents = loader.load()
Each loader handles the specifics of its file format, returning a consistent Document structure.
TEXT SPLITTERS: CHUNKING DOCUMENTS
Large documents need to be split into smaller chunks for effective retrieval. LangChain provides sophisticated text splitters.
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Create a text splitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=200,
chunk_overlap=50,
length_function=len,
separators=["\n\n", "\n", " ", ""]
)
# Split the documents
split_docs = text_splitter.split_documents(documents)
print(f"Original documents: {len(documents)}")
print(f"Split chunks: {len(split_docs)}")
print(f"\nFirst chunk:")
print(split_docs[0].page_content)
print(f"Metadata: {split_docs[0].metadata}")
The RecursiveCharacterTextSplitter tries to split on paragraph boundaries first, then sentences, then words, then characters. This preserves semantic coherence better than naive splitting.
The chunk_overlap parameter ensures that context at chunk boundaries isn't lost. If a sentence is split across chunks, the overlap captures it in both chunks.
Let's see how different splitters work:
from langchain.text_splitter import CharacterTextSplitter
# Simple character splitter
simple_splitter = CharacterTextSplitter(
chunk_size=200,
chunk_overlap=0,
separator=" "
)
simple_chunks = simple_splitter.split_documents(documents)
print(f"Recursive splitter: {len(split_docs)} chunks")
print(f"Simple splitter: {len(simple_chunks)} chunks")
# Compare first chunks
print(f"\nRecursive first chunk:\n{split_docs[0].page_content}\n")
print(f"Simple first chunk:\n{simple_chunks[0].page_content}")
The recursive splitter generally produces more coherent chunks because it respects document structure.
For code, there's a specialized splitter:
from langchain.text_splitter import Language, RecursiveCharacterTextSplitter
# Python code splitter
python_splitter = RecursiveCharacterTextSplitter.from_language(
language=Language.PYTHON,
chunk_size=500,
chunk_overlap=50
)
python_code = """
def fibonacci(n): '''Calculate the nth Fibonacci number.''' if n <= 1: return n return fibonacci(n-1) + fibonacci(n-2)
class Calculator: '''A simple calculator class.'''
def add(self, a, b):
return a + b
def multiply(self, a, b):
return a * b
"""
code_chunks = python_splitter.split_text(python_code)
for i, chunk in enumerate(code_chunks):
print(f"Chunk {i+1}:\n{chunk}\n")
The code splitter understands programming language structure and tries to keep functions and classes together.
VECTOR STORES: STORING AND SEARCHING EMBEDDINGS
Now we need to convert our chunks into embeddings and store them for retrieval. LangChain abstracts the vector store interface.
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
# Create an embedding model
embeddings = HuggingFaceEmbeddings(
model_name="all-MiniLM-L6-v2"
)
# Create a vector store from documents
vectorstore = FAISS.from_documents(
documents=split_docs,
embedding=embeddings
)
print("Vector store created")
print(f"Number of vectors: {vectorstore.index.ntotal}")
The from_documents method handles everything: generating embeddings for each chunk and adding them to the FAISS index. The result is a searchable vector store.
Let's search the vector store:
# Search for similar documents
query = "What are chains in LangChain?"
results = vectorstore.similarity_search(query, k=2)
print(f"Query: {query}\n")
for i, doc in enumerate(results, 1):
print(f"Result {i}:")
print(f"Content: {doc.page_content}")
print(f"Metadata: {doc.metadata}\n")
The similarity_search method finds the most relevant chunks. It returns Document objects with both content and metadata.
You can also get similarity scores:
# Search with scores
results_with_scores = vectorstore.similarity_search_with_score(query, k=3)
print(f"Query: {query}\n")
for doc, score in results_with_scores:
print(f"Score: {score:.4f}")
print(f"Content: {doc.page_content[:100]}...\n")
Lower scores indicate higher similarity (because FAISS uses L2 distance).
The vector store can be saved and loaded:
# Save the vector store
vectorstore.save_local("langchain_vectorstore")
# Load it later
loaded_vectorstore = FAISS.load_local(
"langchain_vectorstore",
embeddings,
allow_dangerous_deserialization=True
)
print("Vector store loaded successfully")
This allows you to build the index once and reuse it across sessions.
LangChain supports many vector store backends:
# Chroma (requires chromadb)
# from langchain_community.vectorstores import Chroma
# vectorstore = Chroma.from_documents(documents=split_docs, embedding=embeddings)
# Pinecone (requires pinecone-client and API key)
# from langchain_community.vectorstores import Pinecone
# vectorstore = Pinecone.from_documents(documents=split_docs, embedding=embeddings, index_name="my-index")
# Qdrant (requires qdrant-client)
# from langchain_community.vectorstores import Qdrant
# vectorstore = Qdrant.from_documents(documents=split_docs, embedding=embeddings, location=":memory:")
The interface remains the same regardless of the backend, making it easy to switch between vector stores.
RETRIEVERS: FLEXIBLE DOCUMENT RETRIEVAL
Vector stores provide similarity search, but retrievers add additional functionality like filtering and hybrid search.
# Convert vector store to retriever
retriever = vectorstore.as_retriever(
search_type="similarity",
search_kwargs={"k": 2}
)
# Use the retriever
query = "How does memory work in LangChain?"
docs = retriever.invoke(query)
print(f"Query: {query}\n")
for doc in docs:
print(f"Content: {doc.page_content}\n")
The as_retriever method creates a retriever from a vector store. Retrievers have a standard interface that works with LangChain's RAG chains.
You can configure different search types:
# Maximum Marginal Relevance (MMR) - balances relevance and diversity
mmr_retriever = vectorstore.as_retriever(
search_type="mmr",
search_kwargs={"k": 3, "fetch_k": 10}
)
docs = mmr_retriever.invoke("Tell me about LangChain features")
print("MMR Results:")
for i, doc in enumerate(docs, 1):
print(f"{i}. {doc.page_content[:80]}...")
MMR retrieves more candidates (fetch_k) and then selects a diverse subset (k) that balances relevance and diversity. This prevents returning multiple very similar chunks.
You can also create custom retrievers with filtering:
from langchain_core.retrievers import BaseRetriever
from langchain_core.documents import Document
from typing import List
class MetadataFilterRetriever(BaseRetriever):
"""Retriever that filters by metadata."""
vectorstore: FAISS
metadata_filter: dict
k: int = 3
def _get_relevant_documents(self, query: str) -> List[Document]:
"""Retrieve documents matching the query and metadata filter."""
# Get more candidates
candidates = self.vectorstore.similarity_search(query, k=self.k * 3)
# Filter by metadata
filtered = [
doc for doc in candidates
if all(doc.metadata.get(key) == value
for key, value in self.metadata_filter.items())
]
# Return top k
return filtered[:self.k]
# Use the custom retriever
filtered_retriever = MetadataFilterRetriever(
vectorstore=vectorstore,
metadata_filter={"topic": "concepts"},
k=2
)
docs = filtered_retriever.invoke("Explain chains")
print("Filtered results (only 'concepts' topic):")
for doc in docs:
print(f"Topic: {doc.metadata['topic']}")
print(f"Content: {doc.page_content[:100]}...\n")
This custom retriever first retrieves candidates, then filters by metadata, and finally returns the top results. This pattern is useful when you want to restrict retrieval to specific document types or sources.
BUILDING RAG WITH LANGCHAIN
Now we can build a complete RAG system. LangChain provides high-level chains that handle the entire RAG workflow.
from langchain.chains import RetrievalQA
from langchain_openai import ChatOpenAI
# Create components
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
retriever = vectorstore.as_retriever(search_kwargs={"k": 2})
# Create RAG chain
qa_chain = RetrievalQA.from_chain_type(
llm=llm,
chain_type="stuff",
retriever=retriever,
return_source_documents=True
)
# Ask questions
query = "What are the different types of memory in LangChain?"
result = qa_chain.invoke({"query": query})
print(f"Question: {query}\n")
print(f"Answer: {result['result']}\n")
print("Source documents:")
for i, doc in enumerate(result['source_documents'], 1):
print(f"{i}. {doc.page_content[:100]}...")
The RetrievalQA chain handles everything: retrieving relevant documents, formatting them into a prompt, calling the LLM, and returning the answer. The return_source_documents flag includes the retrieved chunks in the result.
The chain_type parameter controls how documents are combined. The "stuff" type puts all documents into a single prompt. Let's explore other types:
# Map-reduce: Process each document separately, then combine
mapreduce_chain = RetrievalQA.from_chain_type(
llm=llm,
chain_type="map_reduce",
retriever=retriever
)
# Refine: Iteratively refine the answer with each document
refine_chain = RetrievalQA.from_chain_type(
llm=llm,
chain_type="refine",
retriever=retriever
)
The "map_reduce" type processes each document independently and then combines the results. This is useful for long documents that don't fit in a single prompt. The "refine" type starts with an initial answer and refines it with each additional document.
For more control, use the modern LCEL (LangChain Expression Language) approach:
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
# Define the prompt template
template = """Answer the question based only on the following context:
{context}
Question: {question}
Answer:"""
prompt = ChatPromptTemplate.from_template(template)
# Create the RAG chain
def format_docs(docs):
return "\n\n".join(doc.page_content for doc in docs)
rag_chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
# Use the chain
answer = rag_chain.invoke("How do agents work in LangChain?")
print(answer)
This LCEL chain is more explicit about what happens at each step. The retriever gets relevant documents, format_docs combines them into a string, the prompt template creates the final prompt, the LLM generates an answer, and the output parser extracts the text.
Let's add conversation memory to our RAG system:
from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
# Create memory
memory = ConversationBufferMemory(
memory_key="chat_history",
return_messages=True,
output_key="answer"
)
# Create conversational RAG chain
conversational_chain = ConversationalRetrievalChain.from_llm(
llm=llm,
retriever=retriever,
memory=memory,
return_source_documents=True
)
# Have a conversation
result1 = conversational_chain.invoke({"question": "What is LangChain?"})
print(f"Q: What is LangChain?")
print(f"A: {result1['answer']}\n")
result2 = conversational_chain.invoke({"question": "When was it created?"})
print(f"Q: When was it created?")
print(f"A: {result2['answer']}\n")
result3 = conversational_chain.invoke({"question": "Who created it?"})
print(f"Q: Who created it?")
print(f"A: {result3['answer']}")
The ConversationalRetrievalChain maintains conversation history and uses it to reformulate queries. When you ask "When was it created?", the chain understands that "it" refers to LangChain from the previous question.
AGENTS: AUTONOMOUS REASONING AND TOOL USE
Agents represent LangChain's most advanced capability: giving LLMs the ability to use tools and make decisions autonomously.
from langchain.agents import AgentExecutor, create_react_agent
from langchain.tools import Tool
from langchain_core.prompts import PromptTemplate
# Define some simple tools
def calculator(expression: str) -> str:
"""Evaluate a mathematical expression."""
try:
result = eval(expression)
return f"The result is {result}"
except Exception as e:
return f"Error: {str(e)}"
def word_counter(text: str) -> str:
"""Count the number of words in text."""
count = len(text.split())
return f"The text contains {count} words"
# Create tool objects
tools = [
Tool(
name="Calculator",
func=calculator,
description="Useful for mathematical calculations. Input should be a valid Python expression."
),
Tool(
name="WordCounter",
func=word_counter,
description="Counts the number of words in a text. Input should be the text to count."
)
]
# Create the agent prompt
agent_prompt = PromptTemplate.from_template(
"""Answer the following questions as best you can. You have access to the following tools:
{tools}
Use the following format:
Question: the input question you must answer Thought: you should always think about what to do Action: the action to take, should be one of [{tool_names}] Action Input: the input to the action Observation: the result of the action ... (this Thought/Action/Action Input/Observation can repeat N times) Thought: I now know the final answer Final Answer: the final answer to the original input question
Begin!
Question: {input} Thought: {agent_scratchpad}""" )
# Create the agent
agent = create_react_agent(llm, tools, agent_prompt)
# Create agent executor
agent_executor = AgentExecutor(
agent=agent,
tools=tools,
verbose=True,
max_iterations=5
)
# Use the agent
result = agent_executor.invoke({
"input": "What is 25 * 17, and how many words are in this sentence?"
})
print(f"\nFinal Answer: {result['output']}")
The agent follows a reasoning loop called ReAct (Reasoning and Acting). It thinks about what to do, chooses a tool, executes it, observes the result, and repeats until it has the final answer.
The verbose flag shows the agent's thought process. You'll see it reason about which tools to use and how to combine their results.
Let's create a more practical agent with a search tool:
from langchain_community.tools import DuckDuckGoSearchRun
# Create a search tool
search = DuckDuckGoSearchRun()
# Wrap it as a LangChain tool
search_tool = Tool(
name="Search",
func=search.run,
description="Useful for finding current information about topics. Input should be a search query."
)
# Create agent with search capability
search_tools = [search_tool, tools[0]] # Search and calculator
search_agent = create_react_agent(llm, search_tools, agent_prompt)
search_executor = AgentExecutor(
agent=search_agent,
tools=search_tools,
verbose=True
)
# Ask a question requiring search
result = search_executor.invoke({
"input": "What is the current population of Tokyo, and what is that number divided by 1000?"
})
print(f"\nAnswer: {result['output']}")
The agent searches for Tokyo's population, then uses the calculator to divide it. This demonstrates how agents can chain multiple tools together to accomplish complex tasks.
You can create custom tools for any functionality:
from langchain.tools import BaseTool
from typing import Optional
class DocumentSearchTool(BaseTool):
"""Tool for searching the document knowledge base."""
name = "DocumentSearch"
description = "Search the LangChain documentation. Input should be a question about LangChain."
retriever = None
def __init__(self, retriever):
super().__init__()
self.retriever = retriever
def _run(self, query: str) -> str:
"""Search the documents."""
docs = self.retriever.invoke(query)
if not docs:
return "No relevant information found."
result = "Found the following information:\n\n"
for i, doc in enumerate(docs, 1):
result += f"{i}. {doc.page_content}\n\n"
return result
async def _arun(self, query: str) -> str:
"""Async version."""
raise NotImplementedError("Async not implemented")
# Create the tool
doc_search_tool = DocumentSearchTool(retriever=retriever)
# Create agent with document search
knowledge_tools = [doc_search_tool, tools[0]]
knowledge_agent = create_react_agent(llm, knowledge_tools, agent_prompt)
knowledge_executor = AgentExecutor(
agent=knowledge_agent,
tools=knowledge_tools,
verbose=True
)
# Use it
result = knowledge_executor.invoke({
"input": "What types of memory does LangChain support, and how many are there?"
})
print(f"\nAnswer: {result['output']}")
This agent can search your document knowledge base and perform calculations, combining retrieval with reasoning.
ADVANCED AGENT PATTERNS
LangChain supports more sophisticated agent architectures. The OpenAI Functions agent uses function calling for more reliable tool use:
from langchain.agents import create_openai_functions_agent
# Create a more structured prompt
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
functions_prompt = ChatPromptTemplate.from_messages([
("system", "You are a helpful assistant with access to tools."),
("human", "{input}"),
MessagesPlaceholder(variable_name="agent_scratchpad")
])
# Create OpenAI functions agent
functions_agent = create_openai_functions_agent(llm, tools, functions_prompt)
functions_executor = AgentExecutor(
agent=functions_agent,
tools=tools,
verbose=True
)
# Use it
result = functions_executor.invoke({
"input": "Calculate 15 squared, then count the words in 'LangChain makes building AI applications easier'"
})
print(f"\nResult: {result['output']}")
The OpenAI Functions agent is more reliable because it uses the model's built-in function calling capability rather than parsing text output.
You can also create agents with memory:
from langchain.agents import AgentExecutor, create_react_agent
from langchain.memory import ConversationBufferMemory
# Create memory for the agent
agent_memory = ConversationBufferMemory(
memory_key="chat_history",
return_messages=True
)
# Create agent with memory
memory_agent_prompt = PromptTemplate.from_template(
"""Answer questions using available tools. You have access to:
{tools}
Previous conversation: {chat_history}
Current question: {input}
{agent_scratchpad}""" )
# Note: Integrating memory with agents requires careful prompt engineering
# This is a simplified example
This allows agents to remember previous interactions and use that context in decision-making.
PRODUCTION CONSIDERATIONS
When deploying LangChain applications to production, several considerations become important.
First is error handling. LangChain operations can fail for many reasons: API rate limits, network issues, invalid inputs, or model errors.
from langchain.callbacks import get_openai_callback
import time
def robust_chain_invoke(chain, input_data, max_retries=3):
"""Invoke a chain with retry logic and error handling."""
for attempt in range(max_retries):
try:
with get_openai_callback() as cb:
result = chain.invoke(input_data)
# Log token usage
print(f"Tokens used: {cb.total_tokens}")
print(f"Cost: ${cb.total_cost:.4f}")
return result
except Exception as e:
print(f"Attempt {attempt + 1} failed: {str(e)}")
if attempt < max_retries - 1:
# Exponential backoff
wait_time = 2 ** attempt
print(f"Retrying in {wait_time} seconds...")
time.sleep(wait_time)
else:
print("Max retries reached")
raise
# Use the robust invoke
try:
result = robust_chain_invoke(rag_chain, "What is LangChain?")
print(f"Result: {result}")
except Exception as e:
print(f"Failed after retries: {e}")
The get_openai_callback context manager tracks token usage and costs, which is crucial for monitoring production applications.
Second is caching. Repeated queries should use cached results to save time and money:
from langchain.cache import InMemoryCache
from langchain.globals import set_llm_cache
# Enable caching
set_llm_cache(InMemoryCache())
# Now LLM calls are cached
llm = ChatOpenAI(model="gpt-3.5-turbo")
# First call - hits the API
start = time.time()
result1 = llm.invoke("What is the capital of France?")
time1 = time.time() - start
# Second call - uses cache
start = time.time()
result2 = llm.invoke("What is the capital of France?")
time2 = time.time() - start
print(f"First call: {time1:.2f}s")
print(f"Second call: {time2:.2f}s (cached)")
For persistent caching across sessions, use SQLite or Redis:
from langchain.cache import SQLiteCache
# Use SQLite cache
set_llm_cache(SQLiteCache(database_path=".langchain.db"))
Third is monitoring and logging. Production applications need observability:
from langchain.callbacks import StdOutCallbackHandler
from langchain.callbacks.base import BaseCallbackHandler
class CustomCallbackHandler(BaseCallbackHandler):
"""Custom callback for logging."""
def on_llm_start(self, serialized, prompts, **kwargs):
"""Log when LLM starts."""
print(f"[LLM START] Prompts: {len(prompts)}")
def on_llm_end(self, response, **kwargs):
"""Log when LLM ends."""
print(f"[LLM END] Tokens: {response.llm_output.get('token_usage', {})}")
def on_chain_start(self, serialized, inputs, **kwargs):
"""Log when chain starts."""
print(f"[CHAIN START] {serialized.get('name', 'Unknown')}")
def on_chain_end(self, outputs, **kwargs):
"""Log when chain ends."""
print(f"[CHAIN END]")
# Use the callback
callbacks = [CustomCallbackHandler()]
chain = prompt | llm | StrOutputParser()
result = chain.invoke(
{"topic": "machine learning"},
config={"callbacks": callbacks}
)
Callbacks provide hooks into LangChain's execution flow, allowing you to log, monitor, and debug your applications.
Fourth is streaming for better user experience:
from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
# Create streaming LLM
streaming_llm = ChatOpenAI(
model="gpt-3.5-turbo",
streaming=True,
callbacks=[StreamingStdOutCallbackHandler()]
)
# Use in a chain
streaming_chain = prompt | streaming_llm | StrOutputParser()
print("Streaming response:")
result = streaming_chain.invoke({"topic": "quantum computing"})
The response streams token by token, providing immediate feedback to users.
For RAG systems, you can stream both retrieval and generation:
from langchain.callbacks.manager import CallbackManager
class StreamingRAGHandler(BaseCallbackHandler):
"""Handler for streaming RAG responses."""
def on_retriever_end(self, documents, **kwargs):
"""Called when retrieval completes."""
print(f"\n[Retrieved {len(documents)} documents]\n")
def on_llm_new_token(self, token: str, **kwargs):
"""Called for each new token."""
print(token, end="", flush=True)
# Create streaming RAG chain
streaming_rag = ConversationalRetrievalChain.from_llm(
llm=ChatOpenAI(
model="gpt-3.5-turbo",
streaming=True,
callbacks=[StreamingRAGHandler()]
),
retriever=retriever,
memory=ConversationBufferMemory(
memory_key="chat_history",
return_messages=True,
output_key="answer"
)
)
result = streaming_rag.invoke({"question": "What are agents in LangChain?"})
LANGCHAIN VS DIRECT IMPLEMENTATION
When should you use LangChain versus implementing with lower-level libraries like HuggingFace?
LangChain excels when you need rapid prototyping. Building a RAG system with LangChain takes minutes instead of hours. The abstractions handle common patterns, letting you focus on application logic.
LangChain is ideal for standard workflows. If your use case fits LangChain's patterns (chatbots, RAG, agents), you benefit from battle-tested implementations and community support.
LangChain provides ecosystem integration. It connects to hundreds of services: vector databases, LLM providers, document loaders, and tools. This integration is valuable for complex applications.
However, direct implementation offers more control. You can optimize every detail for your specific use case. You avoid the abstraction overhead and potential bugs in the framework.
Direct implementation is better for learning. Understanding how RAG works at a low level makes you a better AI engineer. LangChain can be a black box that hides important details.
Direct implementation may be necessary for custom requirements. If your use case doesn't fit LangChain's patterns, fighting the framework can be harder than building from scratch.
The best approach often combines both. Use LangChain for rapid prototyping and standard components. Drop down to lower-level libraries for custom or performance-critical parts.
Here's an example mixing LangChain with custom code:
from langchain_community.vectorstores import FAISS
from langchain_community.embeddings import HuggingFaceEmbeddings
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Use LangChain for retrieval
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
vectorstore = FAISS.from_documents(split_docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})
# Use HuggingFace directly for generation (more control)
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")
def custom_rag_answer(question):
"""Custom RAG implementation mixing LangChain and HuggingFace."""
# Use LangChain retriever
docs = retriever.invoke(question)
context = "\n\n".join(doc.page_content for doc in docs)
# Custom prompt formatting
prompt = f"Context:\n{context}\n\nQuestion: {question}\n\nAnswer:"
# Use HuggingFace for generation with custom parameters
inputs = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model.generate(
inputs["input_ids"],
max_length=len(inputs["input_ids"][0]) + 100,
temperature=0.7,
top_p=0.9,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
answer = response[len(prompt):].strip()
return {
"answer": answer,
"context": context,
"sources": [doc.metadata for doc in docs]
}
# Use the hybrid approach
result = custom_rag_answer("What is LangChain?")
print(f"Answer: {result['answer']}")
print(f"\nSources: {result['sources']}")
This approach uses LangChain's retriever (well-tested, handles multiple vector stores) but custom generation logic (full control over parameters and prompt formatting).
ADVANCED LANGCHAIN PATTERNS
Let's explore some advanced patterns that showcase LangChain's power.
First is the multi-query retriever, which generates multiple search queries for better recall:
from langchain.retrievers.multi_query import MultiQueryRetriever
# Create multi-query retriever
multi_retriever = MultiQueryRetriever.from_llm(
retriever=vectorstore.as_retriever(),
llm=llm
)
# It generates multiple queries and combines results
docs = multi_retriever.invoke("How does LangChain handle conversations?")
print(f"Retrieved {len(docs)} documents")
for doc in docs:
print(f"- {doc.page_content[:80]}...")
The multi-query retriever uses the LLM to generate alternative phrasings of the query, retrieves documents for each, and combines the results. This improves recall for ambiguous questions.
Second is the contextual compression retriever, which filters retrieved documents:
from langchain.retrievers import ContextualCompressionRetriever
from langchain.retrievers.document_compressors import LLMChainExtractor
# Create compressor
compressor = LLMChainExtractor.from_llm(llm)
# Create compression retriever
compression_retriever = ContextualCompressionRetriever(
base_compressor=compressor,
base_retriever=retriever
)
# Use it
compressed_docs = compression_retriever.invoke("What are chains?")
print("Compressed documents:")
for doc in compressed_docs:
print(f"- {doc.page_content}")
The compressor uses the LLM to extract only the relevant parts of each document, reducing noise and improving answer quality.
Third is the parent document retriever, which retrieves small chunks but provides larger context:
from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Create storage for parent documents
store = InMemoryStore()
# Create splitters for parent and child chunks
parent_splitter = RecursiveCharacterTextSplitter(chunk_size=400)
child_splitter = RecursiveCharacterTextSplitter(chunk_size=100)
# Create parent document retriever
parent_retriever = ParentDocumentRetriever(
vectorstore=vectorstore,
docstore=store,
child_splitter=child_splitter,
parent_splitter=parent_splitter
)
# Add documents
parent_retriever.add_documents(documents)
# Retrieve - returns parent documents even though search uses child chunks
docs = parent_retriever.invoke("Explain memory systems")
print("Retrieved parent documents:")
for doc in docs:
print(f"Length: {len(doc.page_content)} chars")
print(f"Content: {doc.page_content[:100]}...\n")
This retriever searches using small chunks (for precision) but returns larger parent documents (for context). It's useful when you need both precise retrieval and sufficient context.
Fourth is the ensemble retriever, which combines multiple retrieval methods:
from langchain.retrievers import EnsembleRetriever
from langchain_community.retrievers import BM25Retriever
# Create BM25 retriever (keyword-based)
bm25_retriever = BM25Retriever.from_documents(split_docs)
bm25_retriever.k = 2
# Create ensemble combining semantic and keyword search
ensemble_retriever = EnsembleRetriever(
retrievers=[retriever, bm25_retriever],
weights=[0.5, 0.5]
)
# Use it
docs = ensemble_retriever.invoke("LangChain framework features")
print("Ensemble retrieval results:")
for doc in docs:
print(f"- {doc.page_content[:80]}...")
The ensemble retriever combines semantic search (vector similarity) with keyword search (BM25), providing better results than either method alone.
BUILDING A COMPLETE APPLICATION
Let's bring everything together into a complete LangChain application: a conversational RAG chatbot with tool use.
from langchain_openai import ChatOpenAI
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain.chains import ConversationalRetrievalChain
from langchain.memory import ConversationBufferMemory
from langchain.agents import AgentExecutor, create_openai_functions_agent
from langchain.tools import Tool
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.documents import Document
class AdvancedChatbot:
"""A complete LangChain chatbot with RAG and tools."""
def __init__(self, documents):
"""Initialize the chatbot."""
print("Initializing chatbot...")
# Create LLM
self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.7)
# Process documents
print("Processing documents...")
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=200,
chunk_overlap=50
)
self.chunks = text_splitter.split_documents(documents)
# Create vector store
print("Creating vector store...")
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
self.vectorstore = FAISS.from_documents(self.chunks, embeddings)
self.retriever = self.vectorstore.as_retriever(search_kwargs={"k": 3})
# Create memory
self.memory = ConversationBufferMemory(
memory_key="chat_history",
return_messages=True,
output_key="answer"
)
# Create RAG chain
self.rag_chain = ConversationalRetrievalChain.from_llm(
llm=self.llm,
retriever=self.retriever,
memory=self.memory,
return_source_documents=True
)
# Create tools
self.tools = self._create_tools()
# Create agent
self.agent_executor = self._create_agent()
print("Chatbot ready!")
def _create_tools(self):
"""Create tools for the agent."""
def search_docs(query: str) -> str:
"""Search the knowledge base."""
docs = self.retriever.invoke(query)
if not docs:
return "No information found."
return "\n\n".join(doc.page_content for doc in docs)
def calculate(expression: str) -> str:
"""Evaluate a mathematical expression."""
try:
result = eval(expression)
return str(result)
except:
return "Invalid expression"
return [
Tool(
name="DocumentSearch",
func=search_docs,
description="Search the LangChain knowledge base. Use this for questions about LangChain."
),
Tool(
name="Calculator",
func=calculate,
description="Perform calculations. Input should be a mathematical expression."
)
]
def _create_agent(self):
"""Create the agent."""
prompt = ChatPromptTemplate.from_messages([
("system", "You are a helpful assistant with access to tools."),
("human", "{input}"),
MessagesPlaceholder(variable_name="agent_scratchpad")
])
agent = create_openai_functions_agent(self.llm, self.tools, prompt)
return AgentExecutor(agent=agent, tools=self.tools, verbose=True)
def chat(self, message: str, use_agent: bool = False):
"""Chat with the bot."""
if use_agent:
# Use agent for complex queries
result = self.agent_executor.invoke({"input": message})
return result["output"]
else:
# Use RAG chain for document questions
result = self.rag_chain.invoke({"question": message})
return {
"answer": result["answer"],
"sources": [doc.page_content[:100] for doc in result["source_documents"]]
}
def reset(self):
"""Reset conversation memory."""
self.memory.clear()
# Create and use the chatbot
chatbot = AdvancedChatbot(documents)
# Ask RAG questions
print("\n=== RAG Mode ===")
response = chatbot.chat("What is LangChain?")
print(f"Answer: {response['answer']}")
print(f"Sources: {response['sources']}\n")
response = chatbot.chat("When was it created?")
print(f"Answer: {response['answer']}\n")
# Use agent for complex tasks
print("\n=== Agent Mode ===")
response = chatbot.chat(
"Search for information about chains, then calculate 10 * 5",
use_agent=True
)
print(f"Answer: {response}")
This complete application demonstrates LangChain's power: conversational RAG, memory management, tool use, and agent reasoning, all in a clean, maintainable structure.
WHAT YOU'VE MASTERED
Congratulations! You've journeyed from LangChain basics to building sophisticated AI applications. Let's review what you now understand.
You learned LangChain's core philosophy: composability, abstraction, standardization, and extensibility. You understand how these principles guide the framework's design.
You mastered the fundamental components. You know how to use language models with consistent interfaces. You understand prompt templates for dynamic prompt engineering. You can create chains that compose operations into workflows.
You explored memory systems for maintaining conversation context. You understand the tradeoffs between different memory types and when to use each.
You learned document processing: loading documents from various sources, splitting them intelligently, and creating searchable vector stores. You understand retrievers and their different search strategies.
You built complete RAG systems using LangChain's high-level chains. You understand how retrieval, context injection, and generation work together.
You discovered agents and tool use. You know how agents reason about which tools to use and how to create custom tools for specific tasks.
You explored advanced patterns: multi-query retrieval, contextual compression, parent document retrieval, and ensemble retrieval. You understand when each pattern is appropriate.
You learned production considerations: error handling, caching, monitoring, and streaming. You know how to build robust, observable applications.
Most importantly, you understand when to use LangChain and when to use lower-level libraries. You can make informed architectural decisions.
YOUR NEXT STEPS
Your LangChain journey continues beyond this tutorial. Here are directions to explore.
Build real applications. Create a chatbot for your company's documentation. Build a research assistant that searches papers and synthesizes findings. Create a code assistant that understands your codebase. Real projects teach lessons tutorials can't.
Explore the LangChain ecosystem. Try LangSmith for debugging and monitoring. Experiment with LangServe for deploying chains as APIs. Use LangGraph for building complex, stateful agents.
Study advanced agent patterns. Learn about plan-and-execute agents, multi-agent systems, and hierarchical agents. These patterns enable more sophisticated autonomous behavior.
Experiment with different LLM providers. Try Anthropic's Claude, Google's PaLM, or open-source models through Ollama. Compare their strengths and weaknesses.
Dive into prompt engineering. Learn techniques like chain-of-thought, tree-of-thought, and self-consistency. Master few-shot learning and instruction tuning.
Contribute to the community. LangChain is open source. Report bugs, suggest features, or contribute code. The community is active and welcoming.
Stay current with research. The field evolves rapidly. Follow papers on arXiv, read the LangChain blog, and join discussions on Discord and forums.
Optimize for production. Learn about model quantization, caching strategies, and deployment architectures. Understand the economics of running LLM applications at scale.
RESOURCES FOR CONTINUED LEARNING
The official LangChain documentation is comprehensive and regularly updated. Start with the conceptual guides to deepen your understanding.
The LangChain cookbook provides practical recipes for common tasks. It's an excellent resource for learning patterns and best practices.
LangChain's YouTube channel has tutorials and talks from the team and community. Visual learning complements written documentation.
The LangChain Discord server is active and helpful. Ask questions, share projects, and learn from others building with LangChain.
Harrison Chase's blog posts and talks provide insights into LangChain's design philosophy and future direction.
Papers like ReAct, Tree of Thoughts, and Reflexion explain the reasoning patterns that agents use. Understanding these papers helps you build better agents.
The Awesome LangChain repository curates tools, tutorials, and projects. It's a great way to discover what's possible.
FINAL REFLECTIONS
LangChain represents a shift in how we build AI applications. Instead of writing procedural code that calls APIs, we compose declarative chains that express intent. Instead of managing state manually, we use memory systems. Instead of hard-coding logic, we create agents that reason.
This abstraction has costs and benefits. The costs include learning curve, abstraction overhead, and reduced control. The benefits include rapid development, battle-tested patterns, and ecosystem integration.
The key is knowing when to use which approach. For standard workflows like chatbots and RAG, LangChain accelerates development dramatically. For custom requirements or performance-critical code, lower-level libraries provide necessary control. The best applications often mix both.
You now have the knowledge to build sophisticated AI applications with LangChain. You understand the abstractions, the patterns, and the tradeoffs. You can create chatbots, RAG systems, and autonomous agents.
But more than specific techniques, you've gained a mental model of how to architect AI applications. You understand composition, abstraction, and orchestration. These concepts transcend any particular framework.
The AI field is advancing rapidly. New models, new techniques, and new frameworks emerge constantly. But the fundamentals you've learned here will serve you well. Understanding retrieval, generation, reasoning, and tool use prepares you for whatever comes next.
Keep building. Keep learning. Keep experimenting. The future of AI applications is being written right now, and you're equipped to be part of it.
Welcome to the world of LangChain development. Now go build something extraordinary.