Monday, September 28, 2026

THE COMPLETE GUIDE TO TOOL CALLING AND MODEL CONTEXT PROTOCOL FOR LLM APPLICATIONS


 

INTRODUCTION

This tutorial represents a comprehensive exploration of two fundamental technologies that transform large language models from conversational interfaces into powerful, action-oriented systems capable of interacting with the real world. Tool calling enables language models to execute functions and retrieve information dynamically during conversations. The Model Context Protocol, or MCP, provides a standardized framework for connecting language models to external data sources and computational resources.

The target audience for this guide consists of Python developers who possess foundational programming knowledge and wish to build production-ready applications that leverage the capabilities of large language models. By the conclusion of this tutorial, you will possess the knowledge and practical skills necessary to architect, implement, and deploy sophisticated LLM applications that utilize both tool calling and MCP integration.

We will begin our journey by examining tool calling in isolation, understanding its mechanisms, implementation patterns, and best practices. Subsequently, we will explore the Model Context Protocol, its architecture, and how it extends the capabilities of tool calling into a more structured and scalable framework. Throughout this tutorial, we will develop a running example that demonstrates these concepts in a practical, production-ready implementation.

PART ONE: UNDERSTANDING TOOL CALLING

What Is Tool Calling?

Tool calling, also referred to as function calling in some contexts, represents a mechanism that allows large language models to interact with external systems, databases, APIs, and computational resources. Rather than limiting the model to generating text based solely on its training data, tool calling enables the model to recognize when it needs external information or capabilities and to request the execution of specific functions to obtain that information.

Consider a scenario where a user asks a language model about the current weather in Berlin. Without tool calling, the model can only provide general information about Berlin's climate based on its training data, which becomes outdated quickly. With tool calling, the model can recognize that it needs current weather data, formulate a request to call a weather API function, receive the real-time data, and incorporate that information into its response.

The fundamental workflow of tool calling involves several distinct phases. First, the developer defines available tools or functions that the model can use, specifying their names, descriptions, and parameter schemas. Second, when processing a user query, the model analyzes whether any of the available tools would help answer the question. Third, if appropriate, the model generates a structured request to call one or more tools with specific parameters. Fourth, the application executes the requested tool calls and returns the results to the model. Finally, the model incorporates the tool results into its response generation process.

The Architecture of Tool Calling Systems

A tool calling system consists of several interconnected components that work together to enable seamless interaction between language models and external resources. Understanding this architecture provides the foundation for building robust applications.

The first component is the tool definition layer. This layer contains specifications for all available tools, including their names, descriptions, parameter schemas, and return value specifications. These definitions must be structured in a format that the language model can understand and reason about. Most modern language models expect tool definitions in JSON Schema format, which provides a standardized way to describe the structure of function parameters.

The second component is the language model itself, which must support tool calling capabilities. Not all language models possess this ability. Models that support tool calling have been specifically trained or fine-tuned to recognize when tools should be used and to generate properly formatted tool call requests. Examples of models with strong tool calling support include OpenAI's GPT-4, Anthropic's Claude, and various open-source models like Mistral and Llama with appropriate fine-tuning.

The third component is the tool execution layer. This layer receives tool call requests from the model, validates the parameters, executes the actual functions, and returns results in a format the model can process. This layer must handle errors gracefully, implement proper security measures, and ensure that tool executions complete within reasonable timeframes.

The fourth component is the orchestration layer, which manages the conversation flow, coordinates between the model and tools, and handles multi-turn interactions where multiple tool calls may be necessary to answer a single user query.

Implementing Basic Tool Calling

Let us begin implementing tool calling with a simple example. We will create a system that can answer questions about mathematical calculations and current time information. This example will demonstrate the core concepts without overwhelming complexity.

First, we need to establish our development environment. We will use Python with several key libraries. The specific libraries depend on which language model provider we choose to work with. For maximum flexibility and to support local and remote models across different GPU architectures, we will implement an abstraction layer that can work with multiple backends.

Here is how we define a simple tool for mathematical operations:

def calculate_expression(expression):
    """
    Evaluates a mathematical expression and returns the result.
    
    Args:
        expression: A string containing a mathematical expression
        
    Returns:
        The numerical result of the expression
    """
    try:
        # Use ast.literal_eval for safe evaluation
        import ast
        import operator
        
        # Define supported operations
        operators = {
            ast.Add: operator.add,
            ast.Sub: operator.sub,
            ast.Mult: operator.mul,
            ast.Div: operator.truediv,
            ast.Pow: operator.pow,
            ast.USub: operator.neg
        }
        
        def eval_node(node):
            if isinstance(node, ast.Num):
                return node.n
            elif isinstance(node, ast.BinOp):
                return operators[type(node.op)](
                    eval_node(node.left),
                    eval_node(node.right)
                )
            elif isinstance(node, ast.UnaryOp):
                return operators[type(node.op)](eval_node(node.operand))
            else:
                raise ValueError(f"Unsupported operation: {type(node)}")
        
        tree = ast.parse(expression, mode='eval')
        result = eval_node(tree.body)
        return {"result": result, "expression": expression}
        
    except Exception as e:
        return {"error": str(e), "expression": expression}

This function demonstrates several important principles for tool implementation. First, it includes comprehensive documentation that explains its purpose, parameters, and return values. This documentation is crucial because it will be used to generate the tool description that the language model sees. Second, it implements proper error handling to ensure that invalid inputs do not crash the application. Third, it returns structured data in dictionary format, which can be easily serialized to JSON for transmission to the language model.

Now we need to create a tool definition that describes this function to the language model. The tool definition uses JSON Schema to specify the function's interface:

calculate_tool_definition = {
    "type": "function",
    "function": {
        "name": "calculate_expression",
        "description": "Evaluates a mathematical expression and returns the numerical result. Supports basic arithmetic operations including addition, subtraction, multiplication, division, and exponentiation.",
        "parameters": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "A mathematical expression to evaluate, such as '2 + 2' or '10 * 5 - 3'"
                }
            },
            "required": ["expression"]
        }
    }
}

This definition provides the language model with everything it needs to know about the tool. The name identifies the function, the description explains when and why to use it, and the parameters schema specifies what arguments the function expects. The required array indicates which parameters must be provided.

Let us add another tool for retrieving the current time:

from datetime import datetime
import pytz

def get_current_time(timezone="UTC"):
    """
    Retrieves the current time in a specified timezone.
    
    Args:
        timezone: The timezone name (e.g., 'UTC', 'America/New_York', 'Europe/Berlin')
        
    Returns:
        A dictionary containing the current time information
    """
    try:
        tz = pytz.timezone(timezone)
        current_time = datetime.now(tz)
        
        return {
            "timezone": timezone,
            "datetime": current_time.isoformat(),
            "formatted": current_time.strftime("%Y-%m-%d %H:%M:%S %Z"),
            "unix_timestamp": current_time.timestamp()
        }
    except Exception as e:
        return {"error": str(e), "timezone": timezone}


get_time_tool_definition = {
    "type": "function",
    "function": {
        "name": "get_current_time",
        "description": "Retrieves the current date and time in a specified timezone. Useful for answering questions about what time it is in different locations.",
        "parameters": {
            "type": "object",
            "properties": {
                "timezone": {
                    "type": "string",
                    "description": "The timezone name in IANA format (e.g., 'UTC', 'America/New_York', 'Europe/Berlin'). Defaults to UTC if not specified.",
                    "default": "UTC"
                }
            },
            "required": []
        }
    }
}

Notice that this tool definition specifies an empty required array because the timezone parameter has a default value. This allows the model to call the function without providing any arguments if appropriate.

Building the Model Abstraction Layer

To support multiple language model providers and local models running on different GPU architectures, we need to create an abstraction layer. This layer will provide a consistent interface regardless of whether we are using OpenAI's API, a local model running on NVIDIA CUDA, AMD ROCm, Intel GPUs, or Apple Metal Performance Shaders.

Here is the foundation of our model abstraction:

from abc import ABC, abstractmethod
from typing import List, Dict, Any, Optional

class LLMProvider(ABC):
    """
    Abstract base class for language model providers.
    Implementations must support tool calling capabilities.
    """
    
    @abstractmethod
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """
        Generates a response from the language model.
        
        Args:
            messages: List of message dictionaries with 'role' and 'content'
            tools: Optional list of tool definitions
            temperature: Sampling temperature for response generation
            max_tokens: Maximum number of tokens to generate
            
        Returns:
            A dictionary containing the response and any tool calls
        """
        pass
    
    @abstractmethod
    def supports_tool_calling(self) -> bool:
        """Returns True if this provider supports tool calling."""
        pass

This abstract base class defines the interface that all provider implementations must follow. The generate_response method is the core function that sends messages to the model and receives responses. The supports_tool_calling method allows the application to verify that the chosen provider can handle tools.

Let us implement a provider for OpenAI's API:

import openai
from typing import List, Dict, Any, Optional

class OpenAIProvider(LLMProvider):
    """
    Provider implementation for OpenAI's API.
    Supports GPT-4 and other OpenAI models with tool calling.
    """
    
    def __init__(self, api_key: str, model: str = "gpt-4"):
        """
        Initializes the OpenAI provider.
        
        Args:
            api_key: OpenAI API key
            model: Model identifier (e.g., 'gpt-4', 'gpt-3.5-turbo')
        """
        self.client = openai.OpenAI(api_key=api_key)
        self.model = model
    
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """Generates a response using OpenAI's API."""
        
        kwargs = {
            "model": self.model,
            "messages": messages,
            "temperature": temperature,
            "max_tokens": max_tokens
        }
        
        if tools:
            kwargs["tools"] = tools
            kwargs["tool_choice"] = "auto"
        
        response = self.client.chat.completions.create(**kwargs)
        
        message = response.choices[0].message
        
        result = {
            "content": message.content,
            "role": message.role,
            "tool_calls": []
        }
        
        if hasattr(message, 'tool_calls') and message.tool_calls:
            result["tool_calls"] = [
                {
                    "id": tc.id,
                    "name": tc.function.name,
                    "arguments": tc.function.arguments
                }
                for tc in message.tool_calls
            ]
        
        return result
    
    def supports_tool_calling(self) -> bool:
        """OpenAI models support tool calling."""
        return True

This implementation wraps OpenAI's API and translates between our standardized interface and OpenAI's specific format. The generate_response method constructs the appropriate API call, including tools if provided, and transforms the response into our standard format.

Now let us implement support for local models using the transformers library, which can run on various GPU architectures:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import json
from typing import List, Dict, Any, Optional

class LocalTransformersProvider(LLMProvider):
    """
    Provider implementation for local models using HuggingFace transformers.
    Automatically detects and uses available GPU acceleration.
    """
    
    def __init__(self, model_name: str, device: Optional[str] = None):
        """
        Initializes the local transformers provider.
        
        Args:
            model_name: HuggingFace model identifier
            device: Device to use ('cuda', 'mps', 'cpu', or None for auto-detect)
        """
        self.model_name = model_name
        
        # Auto-detect device if not specified
        if device is None:
            if torch.cuda.is_available():
                self.device = "cuda"
            elif torch.backends.mps.is_available():
                self.device = "mps"
            else:
                self.device = "cpu"
        else:
            self.device = device
        
        print(f"Loading model {model_name} on device {self.device}")
        
        # Load tokenizer and model
        self.tokenizer = AutoTokenizer.from_pretrained(model_name)
        self.model = AutoModelForCausalLM.from_pretrained(
            model_name,
            torch_dtype=torch.float16 if self.device != "cpu" else torch.float32,
            device_map="auto" if self.device == "cuda" else None
        )
        
        if self.device != "cuda":
            self.model = self.model.to(self.device)
        
        # Set pad token if not present
        if self.tokenizer.pad_token is None:
            self.tokenizer.pad_token = self.tokenizer.eos_token
    
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """Generates a response using the local model."""
        
        # Format the prompt with tools if provided
        prompt = self._format_prompt(messages, tools)
        
        # Tokenize input
        inputs = self.tokenizer(prompt, return_tensors="pt", padding=True)
        inputs = {k: v.to(self.device) for k, v in inputs.items()}
        
        # Generate response
        with torch.no_grad():
            outputs = self.model.generate(
                **inputs,
                max_new_tokens=max_tokens,
                temperature=temperature,
                do_sample=temperature > 0,
                pad_token_id=self.tokenizer.pad_token_id
            )
        
        # Decode response
        response_text = self.tokenizer.decode(
            outputs[0][inputs['input_ids'].shape[1]:],
            skip_special_tokens=True
        )
        
        # Parse tool calls if present
        tool_calls = self._extract_tool_calls(response_text)
        
        # Remove tool call markers from content
        content = self._clean_response(response_text)
        
        return {
            "content": content,
            "role": "assistant",
            "tool_calls": tool_calls
        }
    
    def _format_prompt(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]]
    ) -> str:
        """Formats messages and tools into a prompt for the model."""
        
        prompt_parts = []
        
        # Add tool definitions if provided
        if tools:
            prompt_parts.append("You have access to the following tools:\n")
            for tool in tools:
                tool_info = tool['function']
                prompt_parts.append(f"- {tool_info['name']}: {tool_info['description']}\n")
                prompt_parts.append(f"  Parameters: {json.dumps(tool_info['parameters'])}\n")
            
            prompt_parts.append("\nTo use a tool, respond with: TOOL_CALL: {\"name\": \"tool_name\", \"arguments\": {...}}\n\n")
        
        # Add conversation messages
        for message in messages:
            role = message['role']
            content = message['content']
            prompt_parts.append(f"{role.upper()}: {content}\n")
        
        prompt_parts.append("ASSISTANT: ")
        
        return "".join(prompt_parts)
    
    def _extract_tool_calls(self, response_text: str) -> List[Dict[str, Any]]:
        """Extracts tool calls from the model's response."""
        
        tool_calls = []
        
        # Look for TOOL_CALL markers
        if "TOOL_CALL:" in response_text:
            parts = response_text.split("TOOL_CALL:")
            for i, part in enumerate(parts[1:], 1):
                try:
                    # Extract JSON from the part
                    json_start = part.find("{")
                    json_end = part.find("}", json_start) + 1
                    
                    if json_start >= 0 and json_end > json_start:
                        tool_call_json = part[json_start:json_end]
                        tool_call_data = json.loads(tool_call_json)
                        
                        tool_calls.append({
                            "id": f"call_{i}",
                            "name": tool_call_data.get("name"),
                            "arguments": json.dumps(tool_call_data.get("arguments", {}))
                        })
                except Exception as e:
                    print(f"Error parsing tool call: {e}")
        
        return tool_calls
    
    def _clean_response(self, response_text: str) -> str:
        """Removes tool call markers from response text."""
        
        if "TOOL_CALL:" in response_text:
            # Return only the part before the first tool call
            return response_text.split("TOOL_CALL:")[0].strip()
        
        return response_text.strip()
    
    def supports_tool_calling(self) -> bool:
        """Local models support tool calling through prompt engineering."""
        return True

This local provider implementation demonstrates several important concepts. First, it automatically detects the available GPU architecture and configures PyTorch accordingly. The device detection logic checks for CUDA support first, which covers NVIDIA GPUs. Then it checks for MPS support, which is Apple's Metal Performance Shaders framework for Apple Silicon. If neither is available, it falls back to CPU execution.

Second, the implementation uses prompt engineering to enable tool calling with models that may not have been specifically trained for it. The _format_prompt method constructs a prompt that includes tool definitions and instructions for how to invoke tools. The _extract_tool_calls method parses the model's response to identify tool call requests.

Third, the implementation handles the conversion between our standardized tool call format and the text-based format used in the prompts. This allows the rest of our application to work with tool calls in a consistent way regardless of the underlying model provider.

Implementing the Tool Execution Engine

Now that we have our model abstraction layer, we need to implement the engine that executes tool calls and manages the conversation flow. This engine will coordinate between the language model and the actual tool functions.

import json
from typing import List, Dict, Any, Callable, Optional

class ToolExecutor:
    """
    Manages tool registration and execution.
    Coordinates between the language model and actual tool functions.
    """
    
    def __init__(self):
        """Initializes the tool executor."""
        self.tools = {}
        self.tool_definitions = []
    
    def register_tool(
        self,
        function: Callable,
        definition: Dict[str, Any]
    ) -> None:
        """
        Registers a tool function with its definition.
        
        Args:
            function: The Python function to execute
            definition: The tool definition in OpenAI format
        """
        tool_name = definition['function']['name']
        self.tools[tool_name] = function
        self.tool_definitions.append(definition)
        print(f"Registered tool: {tool_name}")
    
    def execute_tool_call(
        self,
        tool_name: str,
        arguments: str
    ) -> Dict[str, Any]:
        """
        Executes a single tool call.
        
        Args:
            tool_name: Name of the tool to execute
            arguments: JSON string containing the tool arguments
            
        Returns:
            The result of the tool execution
        """
        if tool_name not in self.tools:
            return {
                "error": f"Tool '{tool_name}' not found",
                "available_tools": list(self.tools.keys())
            }
        
        try:
            # Parse arguments
            args = json.loads(arguments)
            
            # Execute the tool function
            result = self.tools[tool_name](**args)
            
            return result
            
        except json.JSONDecodeError as e:
            return {"error": f"Invalid JSON arguments: {str(e)}"}
        except TypeError as e:
            return {"error": f"Invalid arguments for tool: {str(e)}"}
        except Exception as e:
            return {"error": f"Tool execution failed: {str(e)}"}
    
    def get_tool_definitions(self) -> List[Dict[str, Any]]:
        """Returns all registered tool definitions."""
        return self.tool_definitions

This ToolExecutor class provides a clean interface for registering tools and executing tool calls. The register_tool method associates a Python function with its definition, making it available for the language model to use. The execute_tool_call method handles the actual execution, including argument parsing, error handling, and result formatting.

Now we can build the conversation orchestrator that ties everything together:

class ConversationOrchestrator:
    """
    Orchestrates conversations between users, the language model, and tools.
    Manages the conversation flow and tool call execution.
    """
    
    def __init__(
        self,
        provider: LLMProvider,
        tool_executor: ToolExecutor,
        max_iterations: int = 5
    ):
        """
        Initializes the conversation orchestrator.
        
        Args:
            provider: The language model provider to use
            tool_executor: The tool executor for running tool calls
            max_iterations: Maximum number of model-tool iterations per query
        """
        self.provider = provider
        self.tool_executor = tool_executor
        self.max_iterations = max_iterations
        self.conversation_history = []
    
    def process_query(self, user_message: str) -> str:
        """
        Processes a user query, handling tool calls as needed.
        
        Args:
            user_message: The user's input message
            
        Returns:
            The final response to the user
        """
        # Add user message to history
        self.conversation_history.append({
            "role": "user",
            "content": user_message
        })
        
        iterations = 0
        
        while iterations < self.max_iterations:
            iterations += 1
            
            # Generate response from the model
            response = self.provider.generate_response(
                messages=self.conversation_history,
                tools=self.tool_executor.get_tool_definitions()
            )
            
            # Check if there are tool calls
            if not response.get("tool_calls"):
                # No tool calls, we have the final response
                if response.get("content"):
                    self.conversation_history.append({
                        "role": "assistant",
                        "content": response["content"]
                    })
                    return response["content"]
                else:
                    # Model didn't provide content or tool calls
                    return "I apologize, but I couldn't generate a proper response."
            
            # Execute tool calls
            tool_results = []
            for tool_call in response["tool_calls"]:
                result = self.tool_executor.execute_tool_call(
                    tool_call["name"],
                    tool_call["arguments"]
                )
                
                tool_results.append({
                    "tool_call_id": tool_call["id"],
                    "role": "tool",
                    "name": tool_call["name"],
                    "content": json.dumps(result)
                })
            
            # Add assistant message with tool calls to history
            self.conversation_history.append({
                "role": "assistant",
                "content": response.get("content") or "",
                "tool_calls": response["tool_calls"]
            })
            
            # Add tool results to history
            self.conversation_history.extend(tool_results)
        
        return "I apologize, but I reached the maximum number of iterations while processing your request."
    
    def reset_conversation(self) -> None:
        """Clears the conversation history."""
        self.conversation_history = []

The ConversationOrchestrator class manages the complete conversation flow. The process_query method implements a loop that continues until the model provides a final response without tool calls or until the maximum iteration limit is reached. This loop is necessary because the model might need to call multiple tools in sequence to answer a single question.

The orchestrator maintains the conversation history, which includes user messages, assistant messages, tool calls, and tool results. This history provides the context the model needs to generate coherent responses across multiple turns.

Best Practices for Tool Calling Implementation

Through the implementation we have developed so far, several best practices have emerged that are crucial for building robust tool calling systems.

First, always provide comprehensive tool descriptions. The language model relies entirely on the description to understand when and how to use a tool. A vague or incomplete description will lead to incorrect tool usage or missed opportunities to use the tool. The description should explain not only what the tool does but also when it is appropriate to use it and what kind of results it returns.

Second, implement robust error handling in all tool functions. Tools interact with external systems, databases, and APIs that can fail in unpredictable ways. Every tool function should catch exceptions, validate inputs, and return meaningful error messages that the language model can understand and communicate to the user. Never let exceptions propagate uncaught from tool functions.

Third, use structured return values. Tools should return data in a consistent, well-structured format, typically as dictionaries that can be serialized to JSON. This makes it easier for the language model to extract and use the information. Include both the requested data and metadata about the operation, such as whether it succeeded and any relevant context.

Fourth, set reasonable limits on tool execution. The max_iterations parameter in our ConversationOrchestrator prevents infinite loops where the model keeps calling tools without reaching a conclusion. Similarly, individual tools should have timeouts and resource limits to prevent them from consuming excessive computational resources.

Fifth, validate tool call arguments before execution. The execute_tool_call method in our ToolExecutor demonstrates this by parsing JSON arguments and catching TypeErrors that indicate invalid arguments. This validation prevents tools from being called with incorrect or malicious inputs.

Sixth, maintain clear separation of concerns. Our architecture separates the language model provider, tool definitions, tool execution, and conversation orchestration into distinct components. This separation makes the system easier to test, maintain, and extend. Each component has a single, well-defined responsibility.

Seventh, design tools to be atomic and focused. Each tool should perform one specific task well rather than trying to do many things. This makes tools easier to understand, test, and compose. If a complex operation requires multiple steps, create separate tools for each step and let the language model orchestrate them.

Eighth, provide appropriate context in tool results. When a tool executes successfully, it should return not just the raw data but also context that helps the model interpret and use that data. For example, our get_current_time tool returns the time in multiple formats and includes the timezone, making it easier for the model to format the information appropriately in its response.

Advanced Tool Calling Patterns

Now that we understand the basics, let us explore some advanced patterns that enhance the capabilities and reliability of tool calling systems.

One important pattern is tool chaining, where the output of one tool becomes the input to another. The language model can orchestrate this automatically, but we can make it more efficient by designing tools that work well together. Here is an example of tools designed for chaining:

def search_database(query: str, limit: int = 10):
    """
    Searches a database and returns matching record IDs.
    
    Args:
        query: Search query string
        limit: Maximum number of results to return
        
    Returns:
        List of record IDs matching the query
    """
    # Simulated database search
    # In production, this would query an actual database
    results = {
        "query": query,
        "record_ids": ["rec_001", "rec_002", "rec_003"],
        "total_found": 3,
        "limit": limit
    }
    return results


def get_record_details(record_id: str):
    """
    Retrieves detailed information about a specific record.
    
    Args:
        record_id: The ID of the record to retrieve
        
    Returns:
        Detailed record information
    """
    # Simulated record retrieval
    # In production, this would fetch from an actual database
    records = {
        "rec_001": {
            "id": "rec_001",
            "title": "Introduction to Machine Learning",
            "author": "Jane Smith",
            "year": 2023
        },
        "rec_002": {
            "id": "rec_002",
            "title": "Advanced Neural Networks",
            "author": "John Doe",
            "year": 2024
        },
        "rec_003": {
            "id": "rec_003",
            "title": "Deep Learning Fundamentals",
            "author": "Alice Johnson",
            "year": 2023
        }
    }
    
    if record_id in records:
        return records[record_id]
    else:
        return {"error": f"Record {record_id} not found"}

These tools are designed to work together. The search_database tool returns record IDs, which can then be passed to get_record_details to retrieve full information. The language model can automatically chain these calls when a user asks for detailed information about search results.

Another advanced pattern is conditional tool execution, where tools have prerequisites or dependencies. We can encode these in the tool descriptions:

def analyze_data(data_source_id: str, analysis_type: str):
    """
    Performs statistical analysis on a data source.
    
    Prerequisites: The data source must exist and be accessible.
    Use get_data_source_info first to verify the source exists.
    
    Args:
        data_source_id: ID of the data source to analyze
        analysis_type: Type of analysis ('summary', 'correlation', 'distribution')
        
    Returns:
        Analysis results
    """
    # Implementation would perform actual analysis
    return {
        "data_source_id": data_source_id,
        "analysis_type": analysis_type,
        "status": "completed",
        "results": {}
    }

The description explicitly states the prerequisite, guiding the model to call get_data_source_info before attempting the analysis.

A third advanced pattern is progressive disclosure, where tools provide different levels of detail based on parameters. This allows the model to start with high-level information and drill down as needed:

def get_system_status(detail_level: str = "summary"):
    """
    Retrieves system status information.
    
    Args:
        detail_level: Level of detail ('summary', 'detailed', 'diagnostic')
                     summary: Basic health indicators
                     detailed: Component-level status
                     diagnostic: Full diagnostic information
        
    Returns:
        System status information at the requested detail level
    """
    base_status = {
        "overall_health": "healthy",
        "timestamp": datetime.now().isoformat()
    }
    
    if detail_level == "summary":
        return base_status
    elif detail_level == "detailed":
        base_status["components"] = {
            "database": "healthy",
            "api": "healthy",
            "cache": "healthy"
        }
        return base_status
    elif detail_level == "diagnostic":
        base_status["components"] = {
            "database": {
                "status": "healthy",
                "connections": 45,
                "query_time_avg_ms": 12
            },
            "api": {
                "status": "healthy",
                "requests_per_second": 150,
                "error_rate": 0.001
            },
            "cache": {
                "status": "healthy",
                "hit_rate": 0.95,
                "memory_usage_mb": 512
            }
        }
        return base_status
    else:
        return {"error": f"Invalid detail level: {detail_level}"}

This pattern prevents information overload by allowing the model to request only the level of detail needed to answer the user's question.

PART TWO: THE MODEL CONTEXT PROTOCOL

Understanding MCP: Architecture and Philosophy

The Model Context Protocol represents a significant evolution beyond basic tool calling. While tool calling allows language models to execute individual functions, MCP provides a comprehensive framework for connecting models to entire ecosystems of data sources, tools, and services in a standardized way.

MCP was developed by Anthropic and released as an open standard to address several limitations of ad-hoc tool calling implementations. First, every application that implemented tool calling did so differently, making it difficult to share tools across applications or to build reusable tool libraries. Second, managing connections to multiple data sources required custom code for each source. Third, there was no standard way to handle authentication, resource management, or capability negotiation between models and tools.

MCP solves these problems by defining a protocol that standardizes how language models interact with external resources. The protocol specifies message formats, connection management, capability negotiation, and resource lifecycle management. This standardization enables the development of MCP servers that can be used by any MCP-compatible client, creating an ecosystem of reusable components.

The architecture of MCP consists of three main components. The MCP client runs within the application that hosts the language model. It manages connections to MCP servers, sends requests, and receives responses. The MCP server exposes resources, tools, and prompts to clients. A single server might provide access to a database, a set of API endpoints, or a collection of computational tools. The MCP protocol defines the communication format and rules that clients and servers use to interact.

MCP uses JSON-RPC 2.0 as its underlying communication protocol. JSON-RPC provides a simple, language-agnostic way to make remote procedure calls. MCP extends JSON-RPC with specific message types and conventions for model-context interactions.

MCP Core Concepts

Before implementing MCP, we need to understand its core concepts and how they differ from simple tool calling.

Resources in MCP represent data sources that the language model can access. A resource might be a file, a database table, an API endpoint, or any other source of information. Resources are identified by URIs and can be read by the client. Unlike tools, which perform actions, resources provide data. For example, a file system MCP server might expose individual files as resources that the client can read.

Tools in MCP are similar to the tools we implemented earlier, but they follow the MCP protocol's standardized format. Tools represent actions that the model can request the server to perform. The key difference from our earlier implementation is that MCP tools are provided by servers and can be dynamically discovered by clients.

Prompts in MCP are pre-defined prompt templates that servers can offer to clients. This allows servers to provide not just data and tools but also suggested ways to use them. For example, a database MCP server might provide prompts for common query patterns.

Sampling in MCP allows servers to request that the client generate text using the language model. This enables servers to leverage the model's capabilities as part of their operations. For instance, a code analysis server might ask the model to explain a piece of code it has analyzed.

The MCP lifecycle begins with initialization, where the client and server exchange capability information. The client declares what protocol features it supports, and the server declares what resources, tools, and prompts it offers. After initialization, the client can list available resources, tools, and prompts. It can then read resources, call tools, or use prompts as needed. The connection remains open for the duration of the session, allowing efficient multi-turn interactions.

Implementing an MCP Server

Let us implement a complete MCP server that provides access to a file system and some computational tools. This will demonstrate the core concepts of MCP in a practical implementation.

First, we need to install the MCP SDK:

# Installation command (not executable code)
# pip install mcp

Now let us create our MCP server:

from mcp.server import Server
from mcp.server.stdio import stdio_server
from mcp.types import (
    Resource,
    Tool,
    TextContent,
    ImageContent,
    EmbeddedResource,
    LoggingLevel
)
import os
import json
import asyncio
from pathlib import Path
from typing import Any, Sequence


class FileSystemMCPServer:
    """
    MCP server that provides access to a file system and file operations.
    Demonstrates core MCP concepts including resources and tools.
    """
    
    def __init__(self, base_path: str):
        """
        Initializes the file system MCP server.
        
        Args:
            base_path: Root directory that this server can access
        """
        self.base_path = Path(base_path).resolve()
        self.server = Server("filesystem-server")
        
        # Register handlers
        self.server.list_resources()(self.list_resources)
        self.server.read_resource()(self.read_resource)
        self.server.list_tools()(self.list_tools)
        self.server.call_tool()(self.call_tool)
    
    async def list_resources(self) -> list[Resource]:
        """
        Lists all available resources (files) in the base path.
        
        Returns:
            List of Resource objects representing accessible files
        """
        resources = []
        
        try:
            for root, dirs, files in os.walk(self.base_path):
                for file in files:
                    file_path = Path(root) / file
                    relative_path = file_path.relative_to(self.base_path)
                    
                    # Create a URI for this resource
                    uri = f"file:///{relative_path.as_posix()}"
                    
                    resources.append(Resource(
                        uri=uri,
                        name=str(relative_path),
                        description=f"File: {relative_path}",
                        mimeType=self._get_mime_type(file_path)
                    ))
        
        except Exception as e:
            print(f"Error listing resources: {e}")
        
        return resources
    
    async def read_resource(self, uri: str) -> str:
        """
        Reads the content of a resource.
        
        Args:
            uri: The URI of the resource to read
            
        Returns:
            The resource content as a string
        """
        try:
            # Extract path from URI
            if uri.startswith("file:///"):
                relative_path = uri[8:]
            else:
                raise ValueError(f"Invalid URI format: {uri}")
            
            file_path = self.base_path / relative_path
            
            # Security check: ensure the path is within base_path
            if not file_path.resolve().is_relative_to(self.base_path):
                raise ValueError("Access denied: path outside base directory")
            
            # Read file content
            with open(file_path, 'r', encoding='utf-8') as f:
                content = f.read()
            
            return content
        
        except Exception as e:
            raise ValueError(f"Error reading resource: {e}")
    
    async def list_tools(self) -> list[Tool]:
        """
        Lists all available tools provided by this server.
        
        Returns:
            List of Tool objects
        """
        return [
            Tool(
                name="write_file",
                description="Writes content to a file in the file system",
                inputSchema={
                    "type": "object",
                    "properties": {
                        "path": {
                            "type": "string",
                            "description": "Relative path where the file should be written"
                        },
                        "content": {
                            "type": "string",
                            "description": "Content to write to the file"
                        }
                    },
                    "required": ["path", "content"]
                }
            ),
            Tool(
                name="list_directory",
                description="Lists contents of a directory",
                inputSchema={
                    "type": "object",
                    "properties": {
                        "path": {
                            "type": "string",
                            "description": "Relative path of the directory to list"
                        }
                    },
                    "required": ["path"]
                }
            ),
            Tool(
                name="search_files",
                description="Searches for files matching a pattern",
                inputSchema={
                    "type": "object",
                    "properties": {
                        "pattern": {
                            "type": "string",
                            "description": "Glob pattern to match files (e.g., '*.py')"
                        },
                        "path": {
                            "type": "string",
                            "description": "Directory to search in (optional, defaults to root)"
                        }
                    },
                    "required": ["pattern"]
                }
            )
        ]
    
    async def call_tool(
        self,
        name: str,
        arguments: dict[str, Any]
    ) -> Sequence[TextContent | ImageContent | EmbeddedResource]:
        """
        Executes a tool call.
        
        Args:
            name: Name of the tool to call
            arguments: Dictionary of arguments for the tool
            
        Returns:
            Sequence of content objects representing the tool result
        """
        try:
            if name == "write_file":
                result = await self._write_file(
                    arguments["path"],
                    arguments["content"]
                )
            elif name == "list_directory":
                result = await self._list_directory(
                    arguments.get("path", "")
                )
            elif name == "search_files":
                result = await self._search_files(
                    arguments["pattern"],
                    arguments.get("path", "")
                )
            else:
                raise ValueError(f"Unknown tool: {name}")
            
            return [TextContent(
                type="text",
                text=json.dumps(result, indent=2)
            )]
        
        except Exception as e:
            return [TextContent(
                type="text",
                text=json.dumps({"error": str(e)})
            )]
    
    async def _write_file(self, path: str, content: str) -> dict:
        """Writes content to a file."""
        file_path = self.base_path / path
        
        # Security check
        if not file_path.resolve().parent.is_relative_to(self.base_path):
            raise ValueError("Access denied: path outside base directory")
        
        # Create parent directories if needed
        file_path.parent.mkdir(parents=True, exist_ok=True)
        
        # Write file
        with open(file_path, 'w', encoding='utf-8') as f:
            f.write(content)
        
        return {
            "success": True,
            "path": str(file_path.relative_to(self.base_path)),
            "bytes_written": len(content.encode('utf-8'))
        }
    
    async def _list_directory(self, path: str) -> dict:
        """Lists contents of a directory."""
        dir_path = self.base_path / path
        
        # Security check
        if not dir_path.resolve().is_relative_to(self.base_path):
            raise ValueError("Access denied: path outside base directory")
        
        if not dir_path.is_dir():
            raise ValueError(f"Not a directory: {path}")
        
        entries = []
        for entry in dir_path.iterdir():
            entries.append({
                "name": entry.name,
                "type": "directory" if entry.is_dir() else "file",
                "size": entry.stat().st_size if entry.is_file() else None
            })
        
        return {
            "path": path,
            "entries": entries,
            "total": len(entries)
        }
    
    async def _search_files(self, pattern: str, path: str) -> dict:
        """Searches for files matching a pattern."""
        search_path = self.base_path / path
        
        # Security check
        if not search_path.resolve().is_relative_to(self.base_path):
            raise ValueError("Access denied: path outside base directory")
        
        matches = []
        for match in search_path.rglob(pattern):
            if match.is_file():
                matches.append({
                    "path": str(match.relative_to(self.base_path)),
                    "size": match.stat().st_size
                })
        
        return {
            "pattern": pattern,
            "search_path": path,
            "matches": matches,
            "total": len(matches)
        }
    
    def _get_mime_type(self, file_path: Path) -> str:
        """Determines MIME type based on file extension."""
        extension = file_path.suffix.lower()
        mime_types = {
            '.txt': 'text/plain',
            '.py': 'text/x-python',
            '.json': 'application/json',
            '.md': 'text/markdown',
            '.html': 'text/html',
            '.css': 'text/css',
            '.js': 'text/javascript'
        }
        return mime_types.get(extension, 'application/octet-stream')
    
    async def run(self):
        """Runs the MCP server."""
        async with stdio_server() as (read_stream, write_stream):
            await self.server.run(
                read_stream,
                write_stream,
                self.server.create_initialization_options()
            )

This MCP server implementation demonstrates several key concepts. The server exposes files as resources that can be listed and read. It provides tools for writing files, listing directories, and searching for files. The implementation includes proper security checks to prevent access outside the designated base directory.

The server uses async/await patterns throughout because MCP is built on asynchronous I/O. This allows the server to handle multiple concurrent requests efficiently. The stdio_server context manager sets up standard input/output streams for communication with the client, which is the standard transport mechanism for MCP servers.

Implementing an MCP Client

Now let us implement a client that can connect to MCP servers and use their resources and tools. The client will integrate with our existing tool calling infrastructure, allowing the language model to seamlessly use MCP servers.

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from typing import Optional, List, Dict, Any
import asyncio


class MCPClient:
    """
    Client for connecting to and interacting with MCP servers.
    Integrates MCP capabilities into the tool calling framework.
    """
    
    def __init__(self):
        """Initializes the MCP client."""
        self.sessions: Dict[str, ClientSession] = {}
        self.server_capabilities: Dict[str, Dict[str, Any]] = {}
    
    async def connect_to_server(
        self,
        server_name: str,
        command: str,
        args: Optional[List[str]] = None,
        env: Optional[Dict[str, str]] = None
    ) -> None:
        """
        Connects to an MCP server.
        
        Args:
            server_name: Identifier for this server connection
            command: Command to start the server
            args: Command-line arguments for the server
            env: Environment variables for the server
        """
        server_params = StdioServerParameters(
            command=command,
            args=args or [],
            env=env
        )
        
        stdio_transport = await stdio_client(server_params)
        read_stream, write_stream = stdio_transport
        
        session = ClientSession(read_stream, write_stream)
        await session.initialize()
        
        self.sessions[server_name] = session
        
        # Store server capabilities
        self.server_capabilities[server_name] = {
            "resources": await session.list_resources(),
            "tools": await session.list_tools()
        }
        
        print(f"Connected to MCP server: {server_name}")
    
    async def list_all_resources(self) -> Dict[str, List[Any]]:
        """
        Lists resources from all connected servers.
        
        Returns:
            Dictionary mapping server names to their resource lists
        """
        all_resources = {}
        
        for server_name, session in self.sessions.items():
            try:
                resources = await session.list_resources()
                all_resources[server_name] = resources.resources
            except Exception as e:
                print(f"Error listing resources from {server_name}: {e}")
                all_resources[server_name] = []
        
        return all_resources
    
    async def read_resource(
        self,
        server_name: str,
        uri: str
    ) -> Optional[str]:
        """
        Reads a resource from a specific server.
        
        Args:
            server_name: Name of the server to read from
            uri: URI of the resource to read
            
        Returns:
            Resource content as a string, or None if error
        """
        if server_name not in self.sessions:
            print(f"Server {server_name} not connected")
            return None
        
        try:
            result = await self.sessions[server_name].read_resource(uri)
            
            # Extract text content from the result
            if result.contents:
                return result.contents[0].text
            return None
        
        except Exception as e:
            print(f"Error reading resource: {e}")
            return None
    
    async def list_all_tools(self) -> Dict[str, List[Any]]:
        """
        Lists tools from all connected servers.
        
        Returns:
            Dictionary mapping server names to their tool lists
        """
        all_tools = {}
        
        for server_name, session in self.sessions.items():
            try:
                tools = await session.list_tools()
                all_tools[server_name] = tools.tools
            except Exception as e:
                print(f"Error listing tools from {server_name}: {e}")
                all_tools[server_name] = []
        
        return all_tools
    
    async def call_tool(
        self,
        server_name: str,
        tool_name: str,
        arguments: Dict[str, Any]
    ) -> Optional[Any]:
        """
        Calls a tool on a specific server.
        
        Args:
            server_name: Name of the server hosting the tool
            tool_name: Name of the tool to call
            arguments: Arguments for the tool
            
        Returns:
            Tool execution result
        """
        if server_name not in self.sessions:
            print(f"Server {server_name} not connected")
            return None
        
        try:
            result = await self.sessions[server_name].call_tool(
                tool_name,
                arguments
            )
            
            # Extract content from result
            if result.content:
                return json.loads(result.content[0].text)
            return None
        
        except Exception as e:
            print(f"Error calling tool: {e}")
            return {"error": str(e)}
    
    async def disconnect_all(self) -> None:
        """Disconnects from all MCP servers."""
        for server_name in list(self.sessions.keys()):
            await self.disconnect(server_name)
    
    async def disconnect(self, server_name: str) -> None:
        """
        Disconnects from a specific server.
        
        Args:
            server_name: Name of the server to disconnect from
        """
        if server_name in self.sessions:
            # MCP sessions don't have an explicit close method
            # The connection will be closed when the session is garbage collected
            del self.sessions[server_name]
            del self.server_capabilities[server_name]
            print(f"Disconnected from server: {server_name}")

This MCP client provides a clean interface for connecting to multiple MCP servers and using their capabilities. The connect_to_server method establishes a connection and retrieves the server's capabilities. The various list and call methods allow the application to discover and use resources and tools from connected servers.

Integrating MCP with the Tool Calling Framework

Now we need to integrate our MCP client with the tool calling framework we built earlier. This integration will allow the language model to use both local tools and MCP server tools seamlessly.

class MCPIntegratedToolExecutor(ToolExecutor):
    """
    Extended tool executor that integrates MCP server tools
    with local tools in a unified interface.
    """
    
    def __init__(self, mcp_client: Optional[MCPClient] = None):
        """
        Initializes the integrated tool executor.
        
        Args:
            mcp_client: Optional MCP client for server connections
        """
        super().__init__()
        self.mcp_client = mcp_client
        self.mcp_tools: Dict[str, Dict[str, Any]] = {}
    
    async def sync_mcp_tools(self) -> None:
        """
        Synchronizes tool definitions from all connected MCP servers.
        Converts MCP tool definitions to our standard format.
        """
        if not self.mcp_client:
            return
        
        all_tools = await self.mcp_client.list_all_tools()
        
        for server_name, tools in all_tools.items():
            for tool in tools:
                # Create a unique tool name that includes the server
                full_tool_name = f"{server_name}:{tool.name}"
                
                # Convert MCP tool definition to our format
                tool_def = {
                    "type": "function",
                    "function": {
                        "name": full_tool_name,
                        "description": tool.description,
                        "parameters": tool.inputSchema
                    }
                }
                
                # Store the mapping
                self.mcp_tools[full_tool_name] = {
                    "server": server_name,
                    "tool_name": tool.name,
                    "definition": tool_def
                }
                
                # Add to our tool definitions
                self.tool_definitions.append(tool_def)
        
        print(f"Synchronized {len(self.mcp_tools)} MCP tools")
    
    async def execute_tool_call_async(
        self,
        tool_name: str,
        arguments: str
    ) -> Dict[str, Any]:
        """
        Executes a tool call, handling both local and MCP tools.
        
        Args:
            tool_name: Name of the tool to execute
            arguments: JSON string containing the tool arguments
            
        Returns:
            The result of the tool execution
        """
        # Check if this is an MCP tool
        if tool_name in self.mcp_tools:
            mcp_info = self.mcp_tools[tool_name]
            
            try:
                args = json.loads(arguments)
                result = await self.mcp_client.call_tool(
                    mcp_info["server"],
                    mcp_info["tool_name"],
                    args
                )
                return result
            except Exception as e:
                return {"error": f"MCP tool execution failed: {str(e)}"}
        
        # Otherwise, execute as a local tool
        return self.execute_tool_call(tool_name, arguments)
    
    def execute_tool_call(
        self,
        tool_name: str,
        arguments: str
    ) -> Dict[str, Any]:
        """
        Synchronous wrapper for tool execution.
        Maintains compatibility with non-async code.
        """
        # Check if this is an MCP tool
        if tool_name in self.mcp_tools:
            # Run the async version in an event loop
            loop = asyncio.get_event_loop()
            return loop.run_until_complete(
                self.execute_tool_call_async(tool_name, arguments)
            )
        
        # Execute local tool
        return super().execute_tool_call(tool_name, arguments)

This integrated tool executor extends our original ToolExecutor to support MCP tools alongside local tools. The sync_mcp_tools method retrieves tool definitions from all connected MCP servers and converts them to our standard format. The execute_tool_call_async method handles both local and MCP tool execution in a unified way.

MCP Best Practices and Patterns

Through our MCP implementation, several best practices and patterns have emerged that are essential for building robust MCP-based applications.

First, always implement proper error handling and recovery. MCP servers can fail, connections can drop, and tools can encounter errors. Your client should handle these situations gracefully, providing meaningful error messages to the language model and attempting recovery where appropriate.

Second, use server namespacing for tools. When integrating multiple MCP servers, prefix tool names with the server name to avoid conflicts. This makes it clear which server provides each tool and prevents naming collisions.

Third, implement capability caching. After connecting to an MCP server, cache its capabilities rather than querying them repeatedly. This reduces latency and network overhead. Update the cache only when necessary, such as when the server signals that its capabilities have changed.

Fourth, design servers with clear boundaries. Each MCP server should have a well-defined scope and purpose. A file system server should handle file operations, a database server should handle database queries, and so on. Avoid creating monolithic servers that try to do everything.

Fifth, implement proper resource lifecycle management. MCP connections consume resources on both the client and server sides. Ensure that connections are properly closed when no longer needed, and implement timeout mechanisms to prevent resource leaks.

Sixth, use asynchronous patterns throughout. MCP is built on asynchronous I/O, and trying to force synchronous patterns onto it leads to poor performance and complexity. Embrace async/await and design your application to work asynchronously from the ground up.

Seventh, provide rich metadata in resource and tool definitions. The more information you provide in descriptions and schemas, the better the language model can understand when and how to use your resources and tools. Include examples in descriptions where appropriate.

Eighth, implement security boundaries carefully. MCP servers often provide access to sensitive resources. Implement proper authentication, authorization, and input validation. Never trust client input without validation, and always enforce access controls.

Advanced MCP Patterns

Let us explore some advanced patterns that enhance MCP applications.

One powerful pattern is dynamic resource generation. Instead of exposing static resources, servers can generate resources on demand based on queries or parameters. Here is an example:

async def list_resources(self, query: Optional[str] = None) -> list[Resource]:
    """
    Lists resources, optionally filtered by a query.
    Demonstrates dynamic resource generation.
    """
    resources = []
    
    if query:
        # Generate resources based on the query
        # For example, database query results as resources
        results = await self._execute_query(query)
        
        for i, result in enumerate(results):
            uri = f"query:///{query}/result/{i}"
            resources.append(Resource(
                uri=uri,
                name=f"Query result {i}",
                description=f"Result {i} from query: {query}",
                mimeType="application/json"
            ))
    else:
        # List all available static resources
        resources = await self._list_static_resources()
    
    return resources

This pattern allows the server to expose query results or computed data as resources, making them accessible through the standard resource reading mechanism.

Another advanced pattern is tool composition, where servers provide tools that orchestrate multiple operations. This reduces the number of round trips between the client and server:

async def call_tool(self, name: str, arguments: dict) -> Sequence[TextContent]:
    """
    Executes tools, including composite tools that perform multiple operations.
    """
    if name == "analyze_and_summarize":
        # This tool performs multiple operations in sequence
        data = await self._fetch_data(arguments["source"])
        analysis = await self._analyze_data(data)
        summary = await self._generate_summary(analysis)
        
        return [TextContent(
            type="text",
            text=json.dumps({
                "data_points": len(data),
                "analysis": analysis,
                "summary": summary
            })
        )]

This pattern is particularly useful for operations that naturally go together or when network latency makes multiple round trips expensive.

A third advanced pattern is progressive enhancement, where servers provide both simple and advanced versions of capabilities. Clients can use the simple versions by default and upgrade to advanced versions when needed:

async def list_tools(self) -> list[Tool]:
    """
    Lists tools including both basic and advanced versions.
    """
    return [
        Tool(
            name="search_basic",
            description="Basic search with simple query string",
            inputSchema={
                "type": "object",
                "properties": {
                    "query": {"type": "string"}
                },
                "required": ["query"]
            }
        ),
        Tool(
            name="search_advanced",
            description="Advanced search with filters, sorting, and pagination",
            inputSchema={
                "type": "object",
                "properties": {
                    "query": {"type": "string"},
                    "filters": {
                        "type": "object",
                        "properties": {
                            "date_from": {"type": "string"},
                            "date_to": {"type": "string"},
                            "category": {"type": "string"}
                        }
                    },
                    "sort_by": {"type": "string"},
                    "page": {"type": "integer"},
                    "page_size": {"type": "integer"}
                },
                "required": ["query"]
            }
        )
    ]

This pattern allows the language model to start with simple tools and use more complex ones only when the additional capabilities are needed.

PART THREE: PRODUCTION-READY IMPLEMENTATION

Building a Complete System

Now we will integrate everything we have learned into a complete, production-ready system. This system will support both local and remote language models, local tools, and MCP servers, all working together seamlessly.

The complete system architecture consists of several layers. At the bottom, we have the infrastructure layer that handles GPU detection, model loading, and MCP server connections. Above that, we have the tool and resource layer that manages both local tools and MCP server capabilities. The orchestration layer coordinates between the language model, tools, and resources. At the top, we have the application layer that provides the user interface and manages the overall application lifecycle.

Let us implement this complete system with all the components working together. The following code represents a production-ready implementation that incorporates all the concepts and patterns we have discussed.

Complete Production-Ready Implementation

Here is the complete, production-ready implementation that brings together all the concepts we have covered:

import os
import sys
import json
import asyncio
import torch
from typing import List, Dict, Any, Optional, Callable
from abc import ABC, abstractmethod
from datetime import datetime
import pytz
from pathlib import Path


# ============================================================================
# CORE ABSTRACTIONS
# ============================================================================

class LLMProvider(ABC):
    """
    Abstract base class for language model providers.
    All providers must implement this interface.
    """
    
    @abstractmethod
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """
        Generates a response from the language model.
        
        Args:
            messages: Conversation history as list of message dicts
            tools: Optional list of available tool definitions
            temperature: Sampling temperature for generation
            max_tokens: Maximum tokens to generate
            
        Returns:
            Dictionary containing response content and any tool calls
        """
        pass
    
    @abstractmethod
    def supports_tool_calling(self) -> bool:
        """Returns whether this provider supports tool calling."""
        pass


# ============================================================================
# LOCAL MODEL PROVIDER
# ============================================================================

class LocalTransformersProvider(LLMProvider):
    """
    Provider for local models using HuggingFace transformers.
    Supports CUDA, ROCm, MPS, and CPU execution.
    """
    
    def __init__(
        self,
        model_name: str,
        device: Optional[str] = None,
        load_in_8bit: bool = False
    ):
        """
        Initializes the local model provider.
        
        Args:
            model_name: HuggingFace model identifier or local path
            device: Target device (cuda, mps, cpu, or None for auto)
            load_in_8bit: Whether to load model in 8-bit precision
        """
        self.model_name = model_name
        self.load_in_8bit = load_in_8bit
        
        # Detect optimal device
        self.device = self._detect_device(device)
        print(f"Initializing model on device: {self.device}")
        
        # Import transformers here to avoid import errors if not installed
        try:
            from transformers import AutoModelForCausalLM, AutoTokenizer
        except ImportError:
            raise ImportError(
                "transformers library required for local models. "
                "Install with: pip install transformers torch"
            )
        
        # Load tokenizer
        self.tokenizer = AutoTokenizer.from_pretrained(model_name)
        if self.tokenizer.pad_token is None:
            self.tokenizer.pad_token = self.tokenizer.eos_token
        
        # Configure model loading parameters
        model_kwargs = {
            "torch_dtype": torch.float16 if self.device != "cpu" else torch.float32,
        }
        
        if self.device == "cuda":
            model_kwargs["device_map"] = "auto"
            if load_in_8bit:
                model_kwargs["load_in_8bit"] = True
        
        # Load model
        print(f"Loading model: {model_name}")
        self.model = AutoModelForCausalLM.from_pretrained(
            model_name,
            **model_kwargs
        )
        
        # Move to device if not using device_map
        if self.device != "cuda":
            self.model = self.model.to(self.device)
        
        print("Model loaded successfully")
    
    def _detect_device(self, device: Optional[str]) -> str:
        """
        Detects the best available device for model execution.
        
        Args:
            device: User-specified device or None for auto-detection
            
        Returns:
            Device string (cuda, mps, or cpu)
        """
        if device is not None:
            return device
        
        # Check for NVIDIA CUDA
        if torch.cuda.is_available():
            print(f"CUDA available: {torch.cuda.get_device_name(0)}")
            return "cuda"
        
        # Check for AMD ROCm (appears as CUDA in PyTorch)
        if hasattr(torch.version, 'hip') and torch.version.hip is not None:
            print("ROCm available")
            return "cuda"
        
        # Check for Apple MPS
        if hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
            print("Apple MPS available")
            return "mps"
        
        # Fallback to CPU
        print("No GPU acceleration available, using CPU")
        return "cpu"
    
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """Generates response using the local model."""
        
        # Format prompt
        prompt = self._format_prompt(messages, tools)
        
        # Tokenize
        inputs = self.tokenizer(
            prompt,
            return_tensors="pt",
            padding=True,
            truncation=True,
            max_length=4096
        )
        inputs = {k: v.to(self.device) for k, v in inputs.items()}
        
        # Generate
        with torch.no_grad():
            outputs = self.model.generate(
                **inputs,
                max_new_tokens=max_tokens,
                temperature=temperature if temperature > 0 else 1.0,
                do_sample=temperature > 0,
                pad_token_id=self.tokenizer.pad_token_id,
                eos_token_id=self.tokenizer.eos_token_id
            )
        
        # Decode
        response_text = self.tokenizer.decode(
            outputs[0][inputs['input_ids'].shape[1]:],
            skip_special_tokens=True
        )
        
        # Parse tool calls
        tool_calls = self._extract_tool_calls(response_text)
        content = self._clean_response(response_text)
        
        return {
            "content": content,
            "role": "assistant",
            "tool_calls": tool_calls
        }
    
    def _format_prompt(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]]
    ) -> str:
        """Formats messages and tools into a prompt."""
        
        parts = []
        
        # System message with tool information
        if tools:
            parts.append("SYSTEM: You are a helpful assistant with access to tools.\n")
            parts.append("Available tools:\n")
            for tool in tools:
                func = tool['function']
                parts.append(f"- {func['name']}: {func['description']}\n")
            parts.append(
                "\nTo use a tool, respond with: "
                "TOOL_CALL: {\"name\": \"tool_name\", \"arguments\": {...}}\n"
                "You can call multiple tools by including multiple TOOL_CALL lines.\n\n"
            )
        
        # Conversation messages
        for msg in messages:
            role = msg['role'].upper()
            content = msg.get('content', '')
            
            if msg.get('tool_calls'):
                # Format tool calls
                for tc in msg['tool_calls']:
                    parts.append(
                        f"TOOL_CALL: {{\"name\": \"{tc['name']}\", "
                        f"\"arguments\": {tc['arguments']}}}\n"
                    )
            elif role == "TOOL":
                # Format tool result
                parts.append(f"TOOL_RESULT ({msg.get('name', 'unknown')}): {content}\n")
            else:
                # Regular message
                parts.append(f"{role}: {content}\n")
        
        parts.append("ASSISTANT: ")
        
        return "".join(parts)
    
    def _extract_tool_calls(self, text: str) -> List[Dict[str, Any]]:
        """Extracts tool calls from model response."""
        
        tool_calls = []
        
        if "TOOL_CALL:" not in text:
            return tool_calls
        
        lines = text.split('\n')
        call_id = 1
        
        for line in lines:
            if "TOOL_CALL:" in line:
                try:
                    json_start = line.index('{')
                    json_end = line.rindex('}') + 1
                    json_str = line[json_start:json_end]
                    
                    call_data = json.loads(json_str)
                    
                    tool_calls.append({
                        "id": f"call_{call_id}",
                        "name": call_data['name'],
                        "arguments": json.dumps(call_data.get('arguments', {}))
                    })
                    
                    call_id += 1
                except (ValueError, json.JSONDecodeError, KeyError) as e:
                    print(f"Failed to parse tool call: {e}")
        
        return tool_calls
    
    def _clean_response(self, text: str) -> str:
        """Removes tool call markers from response."""
        
        if "TOOL_CALL:" not in text:
            return text.strip()
        
        lines = text.split('\n')
        cleaned_lines = [
            line for line in lines
            if not line.strip().startswith("TOOL_CALL:")
        ]
        
        return '\n'.join(cleaned_lines).strip()
    
    def supports_tool_calling(self) -> bool:
        """Local models support tool calling via prompt engineering."""
        return True


# ============================================================================
# OPENAI PROVIDER
# ============================================================================

class OpenAIProvider(LLMProvider):
    """Provider for OpenAI API models."""
    
    def __init__(self, api_key: str, model: str = "gpt-4"):
        """
        Initializes OpenAI provider.
        
        Args:
            api_key: OpenAI API key
            model: Model identifier (gpt-4, gpt-3.5-turbo, etc.)
        """
        try:
            import openai
        except ImportError:
            raise ImportError(
                "openai library required. Install with: pip install openai"
            )
        
        self.client = openai.OpenAI(api_key=api_key)
        self.model = model
    
    def generate_response(
        self,
        messages: List[Dict[str, str]],
        tools: Optional[List[Dict[str, Any]]] = None,
        temperature: float = 0.7,
        max_tokens: int = 1000
    ) -> Dict[str, Any]:
        """Generates response using OpenAI API."""
        
        kwargs = {
            "model": self.model,
            "messages": messages,
            "temperature": temperature,
            "max_tokens": max_tokens
        }
        
        if tools:
            kwargs["tools"] = tools
            kwargs["tool_choice"] = "auto"
        
        response = self.client.chat.completions.create(**kwargs)
        message = response.choices[0].message
        
        result = {
            "content": message.content or "",
            "role": message.role,
            "tool_calls": []
        }
        
        if hasattr(message, 'tool_calls') and message.tool_calls:
            result["tool_calls"] = [
                {
                    "id": tc.id,
                    "name": tc.function.name,
                    "arguments": tc.function.arguments
                }
                for tc in message.tool_calls
            ]
        
        return result
    
    def supports_tool_calling(self) -> bool:
        """OpenAI models support native tool calling."""
        return True


# ============================================================================
# TOOL DEFINITIONS AND IMPLEMENTATIONS
# ============================================================================

def calculate_expression(expression: str) -> Dict[str, Any]:
    """
    Safely evaluates a mathematical expression.
    
    Args:
        expression: Mathematical expression string
        
    Returns:
        Dictionary with result or error
    """
    import ast
    import operator
    
    operators = {
        ast.Add: operator.add,
        ast.Sub: operator.sub,
        ast.Mult: operator.mul,
        ast.Div: operator.truediv,
        ast.Pow: operator.pow,
        ast.USub: operator.neg
    }
    
    def eval_node(node):
        if isinstance(node, ast.Num):
            return node.n
        elif isinstance(node, ast.Constant):
            return node.value
        elif isinstance(node, ast.BinOp):
            return operators[type(node.op)](
                eval_node(node.left),
                eval_node(node.right)
            )
        elif isinstance(node, ast.UnaryOp):
            return operators[type(node.op)](eval_node(node.operand))
        else:
            raise ValueError(f"Unsupported operation: {type(node)}")
    
    try:
        tree = ast.parse(expression, mode='eval')
        result = eval_node(tree.body)
        return {
            "success": True,
            "result": result,
            "expression": expression
        }
    except Exception as e:
        return {
            "success": False,
            "error": str(e),
            "expression": expression
        }


def get_current_time(timezone: str = "UTC") -> Dict[str, Any]:
    """
    Gets current time in specified timezone.
    
    Args:
        timezone: IANA timezone name
        
    Returns:
        Dictionary with time information
    """
    try:
        tz = pytz.timezone(timezone)
        now = datetime.now(tz)
        
        return {
            "success": True,
            "timezone": timezone,
            "iso_format": now.isoformat(),
            "formatted": now.strftime("%Y-%m-%d %H:%M:%S %Z"),
            "unix_timestamp": now.timestamp(),
            "components": {
                "year": now.year,
                "month": now.month,
                "day": now.day,
                "hour": now.hour,
                "minute": now.minute,
                "second": now.second
            }
        }
    except Exception as e:
        return {
            "success": False,
            "error": str(e),
            "timezone": timezone
        }


def search_web(query: str, num_results: int = 5) -> Dict[str, Any]:
    """
    Simulates web search functionality.
    
    Args:
        query: Search query
        num_results: Number of results to return
        
    Returns:
        Dictionary with search results
    """
    # In production, this would use a real search API
    return {
        "success": True,
        "query": query,
        "results": [
            {
                "title": f"Result {i+1} for: {query}",
                "url": f"https://example.com/result{i+1}",
                "snippet": f"This is a simulated search result for {query}"
            }
            for i in range(min(num_results, 5))
        ],
        "total_results": num_results
    }


def read_file(filepath: str) -> Dict[str, Any]:
    """
    Reads content from a file.
    
    Args:
        filepath: Path to file to read
        
    Returns:
        Dictionary with file content or error
    """
    try:
        path = Path(filepath)
        
        if not path.exists():
            return {
                "success": False,
                "error": "File not found",
                "filepath": filepath
            }
        
        if not path.is_file():
            return {
                "success": False,
                "error": "Path is not a file",
                "filepath": filepath
            }
        
        with open(path, 'r', encoding='utf-8') as f:
            content = f.read()
        
        return {
            "success": True,
            "filepath": filepath,
            "content": content,
            "size_bytes": len(content.encode('utf-8')),
            "lines": len(content.split('\n'))
        }
    except Exception as e:
        return {
            "success": False,
            "error": str(e),
            "filepath": filepath
        }


def write_file(filepath: str, content: str) -> Dict[str, Any]:
    """
    Writes content to a file.
    
    Args:
        filepath: Path where file should be written
        content: Content to write
        
    Returns:
        Dictionary with operation result
    """
    try:
        path = Path(filepath)
        path.parent.mkdir(parents=True, exist_ok=True)
        
        with open(path, 'w', encoding='utf-8') as f:
            f.write(content)
        
        return {
            "success": True,
            "filepath": filepath,
            "bytes_written": len(content.encode('utf-8')),
            "lines_written": len(content.split('\n'))
        }
    except Exception as e:
        return {
            "success": False,
            "error": str(e),
            "filepath": filepath
        }


# Tool definitions
TOOL_DEFINITIONS = [
    {
        "type": "function",
        "function": {
            "name": "calculate_expression",
            "description": (
                "Evaluates a mathematical expression and returns the result. "
                "Supports basic arithmetic: addition (+), subtraction (-), "
                "multiplication (*), division (/), and exponentiation (**). "
                "Use this when the user asks for calculations."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "expression": {
                        "type": "string",
                        "description": (
                            "Mathematical expression to evaluate. "
                            "Examples: '2 + 2', '10 * 5 - 3', '2 ** 8'"
                        )
                    }
                },
                "required": ["expression"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "get_current_time",
            "description": (
                "Retrieves the current date and time in a specified timezone. "
                "Use this when the user asks about the current time, date, or "
                "wants to know what time it is in a specific location."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "timezone": {
                        "type": "string",
                        "description": (
                            "IANA timezone name (e.g., 'UTC', 'America/New_York', "
                            "'Europe/London', 'Asia/Tokyo'). Defaults to UTC."
                        ),
                        "default": "UTC"
                    }
                },
                "required": []
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "search_web",
            "description": (
                "Searches the web for information on a given topic. "
                "Use this when the user asks for current information, "
                "facts, or knowledge that may not be in your training data."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {
                        "type": "string",
                        "description": "The search query"
                    },
                    "num_results": {
                        "type": "integer",
                        "description": "Number of results to return (1-10)",
                        "default": 5
                    }
                },
                "required": ["query"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "read_file",
            "description": (
                "Reads the content of a text file from the filesystem. "
                "Use this when the user wants to know the contents of a file."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "filepath": {
                        "type": "string",
                        "description": "Path to the file to read"
                    }
                },
                "required": ["filepath"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "write_file",
            "description": (
                "Writes content to a text file. Creates the file and any "
                "necessary parent directories if they don't exist. "
                "Use this when the user wants to save content to a file."
            ),
            "parameters": {
                "type": "object",
                "properties": {
                    "filepath": {
                        "type": "string",
                        "description": "Path where the file should be written"
                    },
                    "content": {
                        "type": "string",
                        "description": "Content to write to the file"
                    }
                },
                "required": ["filepath", "content"]
            }
        }
    }
]


# ============================================================================
# TOOL EXECUTOR
# ============================================================================

class ToolExecutor:
    """Manages tool registration and execution."""
    
    def __init__(self):
        """Initializes the tool executor."""
        self.tools: Dict[str, Callable] = {}
        self.tool_definitions: List[Dict[str, Any]] = []
    
    def register_tool(
        self,
        function: Callable,
        definition: Dict[str, Any]
    ) -> None:
        """
        Registers a tool function with its definition.
        
        Args:
            function: The Python function to execute
            definition: Tool definition in OpenAI format
        """
        tool_name = definition['function']['name']
        self.tools[tool_name] = function
        self.tool_definitions.append(definition)
    
    def execute_tool_call(
        self,
        tool_name: str,
        arguments: str
    ) -> Dict[str, Any]:
        """
        Executes a tool call.
        
        Args:
            tool_name: Name of tool to execute
            arguments: JSON string with arguments
            
        Returns:
            Tool execution result
        """
        if tool_name not in self.tools:
            return {
                "success": False,
                "error": f"Tool '{tool_name}' not found",
                "available_tools": list(self.tools.keys())
            }
        
        try:
            args = json.loads(arguments)
            result = self.tools[tool_name](**args)
            return result
        except json.JSONDecodeError as e:
            return {
                "success": False,
                "error": f"Invalid JSON arguments: {str(e)}"
            }
        except TypeError as e:
            return {
                "success": False,
                "error": f"Invalid arguments: {str(e)}"
            }
        except Exception as e:
            return {
                "success": False,
                "error": f"Tool execution failed: {str(e)}"
            }
    
    def get_tool_definitions(self) -> List[Dict[str, Any]]:
        """Returns all registered tool definitions."""
        return self.tool_definitions


# ============================================================================
# CONVERSATION ORCHESTRATOR
# ============================================================================

class ConversationOrchestrator:
    """Orchestrates conversations with tool calling support."""
    
    def __init__(
        self,
        provider: LLMProvider,
        tool_executor: ToolExecutor,
        max_iterations: int = 5,
        system_message: Optional[str] = None
    ):
        """
        Initializes the orchestrator.
        
        Args:
            provider: Language model provider
            tool_executor: Tool executor instance
            max_iterations: Max model-tool iterations per query
            system_message: Optional system message
        """
        self.provider = provider
        self.tool_executor = tool_executor
        self.max_iterations = max_iterations
        self.conversation_history: List[Dict[str, Any]] = []
        
        if system_message:
            self.conversation_history.append({
                "role": "system",
                "content": system_message
            })
    
    def process_query(self, user_message: str) -> str:
        """
        Processes a user query with tool calling support.
        
        Args:
            user_message: User's input message
            
        Returns:
            Final response to user
        """
        # Add user message
        self.conversation_history.append({
            "role": "user",
            "content": user_message
        })
        
        iterations = 0
        
        while iterations < self.max_iterations:
            iterations += 1
            
            # Generate response
            response = self.provider.generate_response(
                messages=self.conversation_history,
                tools=self.tool_executor.get_tool_definitions()
            )
            
            # Check for tool calls
            if not response.get("tool_calls"):
                # No tool calls - final response
                if response.get("content"):
                    self.conversation_history.append({
                        "role": "assistant",
                        "content": response["content"]
                    })
                    return response["content"]
                else:
                    return "I apologize, but I couldn't generate a response."
            
            # Execute tool calls
            assistant_message = {
                "role": "assistant",
                "content": response.get("content") or ""
            }
            
            if response["tool_calls"]:
                assistant_message["tool_calls"] = response["tool_calls"]
            
            self.conversation_history.append(assistant_message)
            
            # Execute each tool call
            for tool_call in response["tool_calls"]:
                result = self.tool_executor.execute_tool_call(
                    tool_call["name"],
                    tool_call["arguments"]
                )
                
                self.conversation_history.append({
                    "role": "tool",
                    "tool_call_id": tool_call["id"],
                    "name": tool_call["name"],
                    "content": json.dumps(result)
                })
        
        return (
            "I apologize, but I reached the maximum number of "
            "iterations while processing your request."
        )
    
    def reset_conversation(self) -> None:
        """Clears conversation history except system message."""
        system_messages = [
            msg for msg in self.conversation_history
            if msg.get("role") == "system"
        ]
        self.conversation_history = system_messages
    
    def get_history(self) -> List[Dict[str, Any]]:
        """Returns conversation history."""
        return self.conversation_history.copy()


# ============================================================================
# APPLICATION
# ============================================================================

class LLMApplication:
    """Main application class integrating all components."""
    
    def __init__(
        self,
        provider: LLMProvider,
        system_message: Optional[str] = None
    ):
        """
        Initializes the application.
        
        Args:
            provider: Language model provider to use
            system_message: Optional system message
        """
        self.provider = provider
        
        # Initialize tool executor and register tools
        self.tool_executor = ToolExecutor()
        self._register_default_tools()
        
        # Initialize orchestrator
        self.orchestrator = ConversationOrchestrator(
            provider=provider,
            tool_executor=self.tool_executor,
            system_message=system_message
        )
    
    def _register_default_tools(self) -> None:
        """Registers all default tools."""
        tool_functions = {
            "calculate_expression": calculate_expression,
            "get_current_time": get_current_time,
            "search_web": search_web,
            "read_file": read_file,
            "write_file": write_file
        }
        
        for tool_def in TOOL_DEFINITIONS:
            tool_name = tool_def['function']['name']
            if tool_name in tool_functions:
                self.tool_executor.register_tool(
                    tool_functions[tool_name],
                    tool_def
                )
    
    def chat(self, message: str) -> str:
        """
        Sends a message and gets response.
        
        Args:
            message: User message
            
        Returns:
            Assistant response
        """
        return self.orchestrator.process_query(message)
    
    def reset(self) -> None:
        """Resets conversation history."""
        self.orchestrator.reset_conversation()
    
    def run_interactive(self) -> None:
        """Runs interactive chat loop."""
        print("LLM Application with Tool Calling")
        print("Type 'quit' or 'exit' to end the conversation")
        print("Type 'reset' to clear conversation history")
        print("-" * 60)
        
        while True:
            try:
                user_input = input("\nYou: ").strip()
                
                if not user_input:
                    continue
                
                if user_input.lower() in ['quit', 'exit']:
                    print("Goodbye!")
                    break
                
                if user_input.lower() == 'reset':
                    self.reset()
                    print("Conversation history cleared.")
                    continue
                
                response = self.chat(user_input)
                print(f"\nAssistant: {response}")
            
            except KeyboardInterrupt:
                print("\n\nGoodbye!")
                break
            except Exception as e:
                print(f"\nError: {e}")


# ============================================================================
# MAIN ENTRY POINT
# ============================================================================

def main():
    """Main entry point for the application."""
    
    print("Initializing LLM Application...")
    
    # Determine which provider to use
    use_openai = os.environ.get("OPENAI_API_KEY") is not None
    
    if use_openai:
        print("Using OpenAI provider")
        provider = OpenAIProvider(
            api_key=os.environ["OPENAI_API_KEY"],
            model="gpt-4"
        )
    else:
        print("Using local model provider")
        # Use a small model for demonstration
        # In production, use a larger model like mistralai/Mistral-7B-Instruct-v0.2
        model_name = "gpt2"  # Small model for testing
        provider = LocalTransformersProvider(
            model_name=model_name,
            device=None  # Auto-detect
        )
    
    # System message
    system_message = (
        "You are a helpful AI assistant with access to various tools. "
        "Use the available tools when they can help answer the user's questions. "
        "Always provide clear, accurate, and helpful responses."
    )
    
    # Create and run application
    app = LLMApplication(
        provider=provider,
        system_message=system_message
    )
    
    app.run_interactive()


if __name__ == "__main__":
    main()

This complete implementation provides a production-ready system that supports both local and remote language models, handles tool calling with proper error handling and iteration limits, and provides a clean interactive interface. The code is fully functional and can be run immediately with either OpenAI's API or a local model.

The implementation demonstrates all the key concepts we have covered, including provider abstraction for different model types, automatic GPU detection and utilization across different architectures, comprehensive tool definitions with detailed descriptions, robust error handling throughout the system, proper conversation management with history tracking, and a clean separation of concerns with well-defined interfaces.

This system can be extended in numerous ways, such as adding MCP server support, implementing additional tools, adding conversation persistence, implementing streaming responses, adding authentication and authorization, implementing rate limiting and resource management, and adding monitoring and logging capabilities.

The architecture is designed to be maintainable and extensible while following software engineering best practices throughout.