INTRODUCTION: WHY COMBINE PYTHON WITH C AND C++?
Python has become one of the most popular programming languages in the world, beloved for its clean syntax, extensive libraries, and rapid development cycle. However, Python has a well-known limitation: performance. As an interpreted language with dynamic typing, Python cannot match the raw computational speed of compiled languages like C and C++.
This is where the ability to call C and C++ code from Python becomes invaluable. Imagine you are building a machine learning application that needs to perform inference using a large language model. The model itself might be implemented in C++ for maximum speed, but you want to use Python for the application logic, data preprocessing, and user interface. By bridging these two worlds, you get the best of both: Python's ease of use and C++'s performance.
Real-world examples abound. The NumPy library, which provides fast array operations in Python, is largely implemented in C. TensorFlow and PyTorch, the dominant deep learning frameworks, have C++ cores with Python interfaces. Even the Python interpreter itself is written in C. Understanding how to create these bridges opens up powerful possibilities for your own projects.
In this tutorial, we will explore three main approaches to calling C and C++ code from Python. We will start with the simplest method and gradually work our way to more sophisticated techniques. By the end, you will understand not just how to use these tools, but when to use each one and why they work the way they do.
UNDERSTANDING THE FUNDAMENTAL CHALLENGE
Before we dive into specific techniques, let us understand what we are trying to accomplish. Python code runs in the Python interpreter, which is itself a program written in C. When you write a Python function and call it, the interpreter reads your code, converts it to bytecode, and executes it within its own runtime environment.
C and C++ code, on the other hand, compiles directly to machine code. A C function exists as a sequence of processor instructions in memory. There is no interpreter, no runtime environment in the Python sense. The function simply executes when called.
The challenge is this: how do we make Python, running in its interpreter, call a function that exists as raw machine code? And how do we pass data between these two very different environments?
The answer involves several components.
First, we need to compile our C or C++ code into a shared library, also called a dynamic library. On Linux, these files end with the extension .so. On Windows, they end with .dll. On macOS, they end with .dylib. A shared library is a file containing compiled machine code that can be loaded into a running program.
Second, we need a way for Python to load this shared library and locate the functions we want to call. This is where the various tools we will discuss come in. They provide different mechanisms for loading libraries and creating Python-callable wrappers around C and C++ functions.
Third, we need to handle data type conversion. Python integers, strings, and objects are very different from C integers, char arrays, and structs. The bridging tools must convert data back and forth between these representations.
With this conceptual foundation, let us begin with the simplest approach.
METHOD ONE: USING CTYPES TO CALL C FUNCTIONS
The ctypes module is part of Python's standard library, which means you already have it installed. It provides a straightforward way to call functions in shared libraries without writing any additional C code. This makes it perfect for beginners and for simple use cases.
Let us start with the most basic example possible: a C function that adds two integers.
Creating Your First C Function
Create a file named simple.c with the following content:
int add_numbers(int a, int b) {
return a + b;
}
This function is as simple as it gets. It takes two integer parameters and returns their sum. Notice that we have not included any special Python-related code. This is pure C.
Now we need to compile this into a shared library. On Linux or macOS, open a terminal in the directory containing simple.c and run:
gcc -shared -fPIC -o libsimple.so simple.c
On Windows, the command is slightly different:
gcc -shared -o simple.dll simple.c
Let us break down what this command does. The gcc compiler takes our source file simple.c and compiles it. The.shared flag tells gcc to create a shared library rather than a standalone executable. The .fPIC flag, which stands for Position Independent Code, is required on most Unix-like systems for shared libraries. The .o flag specifies the output filename.
After running this command, you should have a file named libsimple.so on Linux or macOS, or simple.dll on Windows. This file contains the compiled machine code for our add_numbers function.
Loading the Library in Python
Now let us write Python code to load this library and call our function. Create a file named use_simple.py:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Call the function
result = lib.add_numbers(5, 3)
print(f"Result: {result}")
When you run this Python script, it should print "Result: 8". Let us examine what happened.
The ctypes.CDLL class is used to load a shared library. CDLL stands for C Dynamic Link Library. When we create a CDLL object with the path to our library file, ctypes loads that library into memory. The library object we get back acts as a gateway to the functions inside.
Once we have the library object, we can access functions by name as attributes. When we write lib.add_numbers, ctypes looks up a function named add_numbers in the loaded library. We can then call this function just like any Python function, passing integer arguments.
Behind the scenes, ctypes is doing several things. It is converting the Python integers 5 and 3 into C integers. It is calling the actual machine code function. And it is converting the C integer result back into a Python integer.
Specifying Argument and Return Types
The previous example worked, but we got lucky. We did not tell ctypes what types of arguments our function expects or what type it returns. For simple cases with integers, ctypes makes reasonable assumptions. But for more complex scenarios, we need to be explicit.
Let us modify our Python code to specify types:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Specify the argument types and return type
lib.add_numbers.argtypes = [ctypes.c_int, ctypes.c_int]
lib.add_numbers.restype = ctypes.c_int
# Call the function
result = lib.add_numbers(5, 3)
print(f"Result: {result}")
The argtypes attribute is set to a list of ctypes type objects. Each element in the list corresponds to one parameter of the C function. Here we specify that both parameters are c_int, which is ctypes' representation of a C integer.
The restype attribute specifies the return type. Again, we use c_int.
Why is this important? First, it makes your code more robust. If you accidentally pass the wrong type of argument, ctypes can catch the error. Second, for some data types, ctypes cannot guess correctly without explicit type information. Third, it serves as documentation, making your code easier to understand.
Working with Floating Point Numbers
Let us extend our example to work with floating point numbers. Add this function to simple.c:
double multiply_floats(double x, double y) {
return x * y;
}
Recompile the library using the same gcc command as before. Now update use_simple.py:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Configure the multiply_floats function
lib.multiply_floats.argtypes = [ctypes.c_double, ctypes.c_double]
lib.multiply_floats.restype = ctypes.c_double
# Call the function
result = lib.multiply_floats(3.14, 2.0)
print(f"Result: {result}")
The pattern is the same, but we use c_double instead of c_int. The c_double type corresponds to the C double type, which is a double-precision floating point number.
This demonstrates an important principle: ctypes provides Python equivalents for all the basic C data types. You have c_int, c_long, c_float, c_double, c_char, c_bool, and many others. Each one knows how to convert between the Python representation and the C representation.
Passing Arrays to C Functions
Arrays are where things get more interesting. In C, an array is essentially a pointer to the first element. When we pass an array from Python to C, we need to handle this pointer concept.
Let us create a C function that sums all elements in an array. Add this to simple.c:
double sum_array(double* arr, int length) {
double total = 0.0;
int i;
for (i = 0; i < length; i++) {
total += arr[i];
}
return total;
}
This function takes a pointer to an array of doubles and the length of the array. It iterates through the array, accumulating the sum.
Recompile the library. Now let us call this function from Python:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Configure the sum_array function
lib.sum_array.argtypes = [ctypes.POINTER(ctypes.c_double), ctypes.c_int]
lib.sum_array.restype = ctypes.c_double
# Create a Python list
numbers = [1.5, 2.5, 3.5, 4.5]
# Convert to a ctypes array
array_type = ctypes.c_double * len(numbers)
c_array = array_type(*numbers)
# Call the function
result = lib.sum_array(c_array, len(numbers))
print(f"Sum: {result}")
Let us unpack this carefully. First, we specify that the first argument is a pointer to double. We do this using ctypes.POINTER(ctypes.c_double). This tells ctypes that the C function expects a memory address pointing to double values.
Next, we create a Python list with some numbers. But we cannot pass this list directly to C. We need to convert it to a C-style array.
The array_type equals ctypes.c_double times len(numbers) creates a new type. This is a ctypes array type that holds a specific number of doubles. Think of it as defining the blueprint for an array.
Then we create an actual array by calling array_type with our numbers as arguments. The asterisk before numbers unpacks the list, passing each element as a separate argument.
Finally, we pass this ctypes array to our C function. The ctypes library automatically converts the array to a pointer when passing it to C.
This example shows how ctypes handles more complex data structures. While it requires more setup than simple integers, the pattern is consistent and logical.
Modifying Arrays In Place
One powerful feature of passing arrays to C is that C can modify the array contents, and those changes are visible in Python. Let us demonstrate this with a function that doubles each element in an array.
Add to simple.c:
void double_array_elements(double* arr, int length) {
int i;
for (i = 0; i < length; i++) {
arr[i] = arr[i] * 2.0;
}
}
This function returns void because it modifies the array in place rather than returning a new value.
Recompile and update the Python code:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Configure the function
lib.double_array_elements.argtypes = [ctypes.POINTER(ctypes.c_double), ctypes.c_int]
lib.double_array_elements.restype = None
# Create and populate array
numbers = [1.0, 2.0, 3.0, 4.0]
array_type = ctypes.c_double * len(numbers)
c_array = array_type(*numbers)
print(f"Before: {list(c_array)}")
# Call the function
lib.double_array_elements(c_array, len(numbers))
print(f"After: {list(c_array)}")
When you run this, you will see that the array values have been doubled. Notice that we set restype to None, which is how we indicate that a C function returns void.
The key insight here is that c_array is not a copy of the Python list. It is a ctypes object that manages a block of memory containing the actual double values. When we pass this to C, we are passing a pointer to that memory. The C function can read and write that memory directly.
This is both powerful and dangerous. It is powerful because it allows efficient data sharing between Python and C without copying large arrays. It is dangerous because C has no bounds checking. If the C code writes beyond the end of the array, it will corrupt memory and likely crash your program.
Working with Strings
Strings present special challenges because Python and C handle them very differently. In Python, strings are immutable Unicode objects. In C, strings are arrays of characters terminated by a null byte.
Let us create a C function that converts a string to uppercase. Add to simple.c:
#include <ctype.h>
void uppercase_string(char* str) {
int i;
for (i = 0; str[i] != '\0'; i++) {
str[i] = toupper(str[i]);
}
}
This function takes a pointer to a character array and modifies it in place, converting each character to uppercase. The loop continues until it encounters the null terminator.
Recompile the library. Now for the Python side:
import ctypes
import sys
# Load the shared library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ctypes.CDLL('./libsimple.so')
else:
lib = ctypes.CDLL('./simple.dll')
# Configure the function
lib.uppercase_string.argtypes = [ctypes.c_char_p]
lib.uppercase_string.restype = None
# Create a mutable buffer
text = b"hello world"
buffer = ctypes.create_string_buffer(text)
print(f"Before: {buffer.value}")
# Call the function
lib.uppercase_string(buffer)
print(f"After: {buffer.value}")
There are several important details here. First, we use c_char_p as the argument type, which represents a pointer to a character array.
Second, we create the string with a b prefix, making it a bytes object rather than a Unicode string. C functions work with bytes, not Unicode.
Third, and most importantly, we use ctypes.create_string_buffer. This creates a mutable buffer containing a copy of our string. We cannot pass a Python bytes object directly because Python strings are immutable. The C function needs to be able to modify the string, so we need a mutable buffer.
The buffer object has a value attribute that gives us the current contents as a bytes object. After calling the C function, we can see that the string has been converted to uppercase.
METHOD TWO: USING CFFI FOR A MORE MODERN APPROACH
While ctypes is convenient and built into Python, it has some limitations. The syntax can be verbose, and it is easy to make mistakes with type declarations. CFFI, which stands for C Foreign Function Interface, is a more modern alternative that addresses these issues.
CFFI is not part of the standard library, so you need to install it first:
pip install cffi
CFFI offers two modes: ABI mode and API mode. ABI mode is similar to ctypes, where you load an existing shared library. API mode is more powerful: it actually compiles C code for you, creating optimized bindings.
Let us start with ABI mode because it is simpler to understand.
CFFI in ABI Mode
We will use the same simple.c file from earlier. Make sure it is compiled into a shared library.
Here is how to call our add_numbers function using CFFI:
from cffi import FFI
import sys
# Create an FFI instance
ffi = FFI()
# Declare the function signature
ffi.cdef("""
int add_numbers(int a, int b);
""")
# Load the library
if sys.platform.startswith('linux') or sys.platform == 'darwin':
lib = ffi.dlopen('./libsimple.so')
else:
lib = ffi.dlopen('./simple.dll')
# Call the function
result = lib.add_numbers(5, 3)
print(f"Result: {result}")
The first step is creating an FFI object. This object will manage all our interactions with C code.
Next, we use the cdef method to declare the function signature. Notice that we write this as actual C syntax, not as Python types. We simply copy the function declaration from the C header file. This is more intuitive than ctypes' approach of specifying argtypes and restype separately.
Then we load the library using dlopen, which is similar to ctypes.CDLL.
Finally, we call the function. The syntax is cleaner than ctypes, and CFFI handles type conversion automatically based on the declaration we provided.
CFFI in API Mode
API mode is where CFFI really shines. Instead of loading a pre-compiled library, CFFI compiles the C code for you and creates optimized bindings. This approach is more robust and often faster.
Let us create a new example. First, create a file named math_ops.c:
double compute_mean(double* values, int count) {
double sum = 0.0;
int i;
for (i = 0; i < count; i++) {
sum += values[i];
}
return sum / count;
}
double compute_variance(double* values, int count, double mean) {
double sum_squared_diff = 0.0;
int i;
for (i = 0; i < count; i++) {
double diff = values[i] - mean;
sum_squared_diff += diff * diff;
}
return sum_squared_diff / count;
}
These functions compute the mean and variance of an array of numbers. Now create a Python script named build_math_ops.py:
from cffi import FFI
ffi = FFI()
# Declare the functions
ffi.cdef("""
double compute_mean(double* values, int count);
double compute_variance(double* values, int count, double mean);
""")
# Set the source code
ffi.set_source("_math_ops",
"""
double compute_mean(double* values, int count) {
double sum = 0.0;
int i;
for (i = 0; i < count; i++) {
sum += values[i];
}
return sum / count;
}
double compute_variance(double* values, int count, double mean) {
double sum_squared_diff = 0.0;
int i;
for (i = 0; i < count; i++) {
double diff = values[i] - mean;
sum_squared_diff += diff * diff;
}
return sum_squared_diff / count;
}
"""
)
# Compile
ffi.compile()
Run this script with python build_math_ops.py. It will create a compiled extension module.
The set_source method is the key. The first argument is the name of the module to create. The second argument is the actual C source code. CFFI will compile this code and create a Python extension module.
Now we can use our compiled module:
from _math_ops import ffi, lib
# Create some test data
values = [1.0, 2.0, 3.0, 4.0, 5.0]
# Convert to C array
c_values = ffi.new("double[]", values)
# Compute mean
mean = lib.compute_mean(c_values, len(values))
print(f"Mean: {mean}")
# Compute variance
variance = lib.compute_variance(c_values, len(values), mean)
print(f"Variance: {variance}")
Notice how we import from the module we created. The ffi object provides utilities for creating C data structures. The lib object contains our C functions.
The ffi.new method creates a C array. We specify the type as a string using C syntax. CFFI allocates memory for the array and initializes it with our Python values.
This approach has several advantages over ABI mode. The bindings are compiled specifically for your platform, making them faster. CFFI can perform more optimizations because it sees the actual C code. And the build process is more reproducible because the C code is embedded in your Python script.
Understanding Memory Management in CFFI
One crucial aspect of CFFI is memory management. When you create C objects using ffi.new, CFFI manages their lifetime. The memory is automatically freed when the Python object is garbage collected.
However, you need to be careful about object lifetimes. Consider this example:
def create_array():
values = [1.0, 2.0, 3.0]
c_array = ffi.new("double[]", values)
return c_array
arr = create_array()
# arr is still valid here
This works correctly. The c_array object is returned from the function, so Python keeps it alive. The memory will not be freed until arr goes out of scope.
But this would be dangerous:
def get_pointer():
values = [1.0, 2.0, 3.0]
c_array = ffi.new("double[]", values)
return ffi.cast("double*", c_array)
ptr = get_pointer()
# ptr points to freed memory - DANGER!
Here we cast the array to a raw pointer and return that. The c_array object is not returned, so it gets garbage collected when the function returns. The pointer now points to freed memory. Using it would cause undefined behavior.
The lesson is to always keep references to CFFI objects for as long as you need the underlying memory.
METHOD THREE: USING PYBIND11 FOR C++ INTEGRATION
So far we have focused on C. But what about C++? C++ has classes, templates, operator overloading, and many other features that do not exist in C. While you can use ctypes or CFFI with C++ by writing C wrapper functions, there is a better way: Pybind11.
Pybind11 is a header-only library that makes it easy to create Python bindings for C++ code. It leverages modern C++ features to provide a clean, intuitive syntax.
To use Pybind11, you need to install it:
pip install pybind11
You also need a C++ compiler. On Linux, you likely have g++. On Windows, you can use Visual Studio or MinGW. On macOS, you can use clang++ from Xcode.
Your First Pybind11 Module
Let us create a simple C++ module. Create a file named example.cpp:
#include <pybind11/pybind11.h>
namespace py = pybind11;
int add(int a, int b) {
return a + b;
}
PYBIND11_MODULE(example, m) {
m.doc() = "A simple example module";
m.def("add", &add, "A function that adds two numbers");
}
The first line includes the Pybind11 header. We then create a namespace alias for convenience.
The add function is a normal C++ function. Nothing special here.
The magic happens in the PYBIND11_MODULE macro. This macro creates the Python module. The first argument is the module name. The second argument is a variable name for the module object.
Inside the macro, we set the module's documentation string. Then we use m.def to define a function. We provide the Python name, a pointer to the C++ function, and a documentation string.
To compile this, we need a more complex command. On Linux:
g++ -O3 -Wall -shared -std=c++11 -fPIC $(python3 -m pybind11 --includes) example.cpp -o example$(python3-config --extension-suffix)
On Windows with Visual Studio, you would use cl.exe with appropriate flags. The details vary by platform, but the key points are: we need to include the Pybind11 headers, compile with C++11 or later, and create a shared library with the right extension.
After compiling, you can use the module in Python:
import example
result = example.add(5, 3)
print(f"Result: {result}")
It is that simple. Pybind11 handles all the type conversion automatically. Python integers are converted to C++ integers, the function is called, and the result is converted back.
Binding C++ Classes
The real power of Pybind11 shows when working with C++ classes. Let us create a simple Vector class. Create vector_module.cpp:
#include <pybind11/pybind11.h>
#include <cmath>
namespace py = pybind11;
class Vector2D {
public:
double x, y;
Vector2D(double x, double y) : x(x), y(y) {}
double length() const {
return std::sqrt(x * x + y * y);
}
Vector2D operator+(const Vector2D& other) const {
return Vector2D(x + other.x, y + other.y);
}
Vector2D operator*(double scalar) const {
return Vector2D(x * scalar, y * scalar);
}
};
PYBIND11_MODULE(vector_module, m) {
py::class_<Vector2D>(m, "Vector2D")
.def(py::init<double, double>())
.def_readwrite("x", &Vector2D::x)
.def_readwrite("y", &Vector2D::y)
.def("length", &Vector2D::length)
.def("__add__", &Vector2D::operator+)
.def("__mul__", &Vector2D::operator*);
}
This is a simple 2D vector class with a constructor, public members, a length method, and operator overloading for addition and scalar multiplication.
The binding code uses py::class_ to create a Python class. The template parameter is the C++ class type. We pass the module object and the Python class name.
The def method with py::init specifies the constructor. The template parameters to py::init are the constructor argument types.
The def_readwrite method exposes member variables. These become Python attributes that can be read and written.
Regular methods are bound with def, just like module-level functions.
For operator overloading, we bind the operators to Python's special methods. The addition operator becomes add, and the multiplication operator becomes mul.
After compiling this module, you can use it naturally in Python:
import vector_module
v1 = vector_module.Vector2D(3.0, 4.0)
v2 = vector_module.Vector2D(1.0, 2.0)
print(f"v1 length: {v1.length()}")
v3 = v1 + v2
print(f"v3: ({v3.x}, {v3.y})")
v4 = v1 * 2.0
print(f"v4: ({v4.x}, {v4.y})")
The C++ class feels like a native Python class. You can create instances, access attributes, call methods, and use operators. Pybind11 handles all the conversions and memory management.
Working with STL Containers
Pybind11 has excellent support for the C++ Standard Template Library. Let us create a function that works with vectors. Create stl_example.cpp:
#include <pybind11/pybind11.h>
#include <pybind11/stl.h>
#include <vector>
#include <numeric>
namespace py = pybind11;
double compute_average(const std::vector<double>& values) {
if (values.empty()) {
return 0.0;
}
double sum = std::accumulate(values.begin(), values.end(), 0.0);
return sum / values.size();
}
std::vector<double> scale_values(const std::vector<double>& values, double factor) {
std::vector<double> result;
result.reserve(values.size());
for (double v : values) {
result.push_back(v * factor);
}
return result;
}
PYBIND11_MODULE(stl_example, m) {
m.def("compute_average", &compute_average);
m.def("scale_values", &scale_values);
}
Notice the second include: pybind11/stl.h. This header provides automatic conversion between Python lists and C++ vectors, Python dictionaries and C++ maps, and other STL containers.
The compute_average function takes a const reference to a vector of doubles and returns their average. The scale_values function takes a vector and a scaling factor, returning a new vector with scaled values.
After compiling, you can use these functions with Python lists:
import stl_example
values = [1.0, 2.0, 3.0, 4.0, 5.0]
avg = stl_example.compute_average(values)
print(f"Average: {avg}")
scaled = stl_example.scale_values(values, 2.0)
print(f"Scaled: {scaled}")
Pybind11 automatically converts the Python list to a C++ vector when calling the function. When the function returns a vector, Pybind11 converts it back to a Python list. This happens transparently, with no extra code required.
This automatic conversion is incredibly convenient, but be aware of the performance implications. Each conversion involves copying all the data. For large containers, this can be expensive. In performance-critical code, you might want to use NumPy arrays instead, which we will discuss shortly.
PRACTICAL EXAMPLE: IMPLEMENTING FAST MATRIX MULTIPLICATION
Now let us put everything together with a realistic example: implementing matrix multiplication in C++ and calling it from Python. This is exactly the kind of task where C++ shines: computationally intensive operations that benefit from compiled code.
We will implement this using Pybind11 because it provides the cleanest interface for C++ code.
The C++ Implementation
Create a file named matrix_ops.cpp:
#include <pybind11/pybind11.h>
#include <pybind11/stl.h>
#include <vector>
#include <stdexcept>
namespace py = pybind11;
// Simple matrix class
class Matrix {
private:
std::vector<double> data;
size_t rows;
size_t cols;
public:
Matrix(size_t rows, size_t cols) : rows(rows), cols(cols) {
data.resize(rows * cols, 0.0);
}
Matrix(size_t rows, size_t cols, const std::vector<double>& values)
: rows(rows), cols(cols), data(values) {
if (values.size() != rows * cols) {
throw std::invalid_argument("Data size does not match dimensions");
}
}
size_t get_rows() const { return rows; }
size_t get_cols() const { return cols; }
double get(size_t i, size_t j) const {
return data[i * cols + j];
}
void set(size_t i, size_t j, double value) {
data[i * cols + j] = value;
}
std::vector<double> get_data() const {
return data;
}
Matrix multiply(const Matrix& other) const {
if (cols != other.rows) {
throw std::invalid_argument("Matrix dimensions incompatible for multiplication");
}
Matrix result(rows, other.cols);
for (size_t i = 0; i < rows; i++) {
for (size_t j = 0; j < other.cols; j++) {
double sum = 0.0;
for (size_t k = 0; k < cols; k++) {
sum += get(i, k) * other.get(k, j);
}
result.set(i, j, sum);
}
}
return result;
}
};
PYBIND11_MODULE(matrix_ops, m) {
m.doc() = "Fast matrix operations implemented in C++";
py::class_<Matrix>(m, "Matrix")
.def(py::init<size_t, size_t>())
.def(py::init<size_t, size_t, const std::vector<double>&>())
.def("get_rows", &Matrix::get_rows)
.def("get_cols", &Matrix::get_cols)
.def("get", &Matrix::get)
.def("set", &Matrix::set)
.def("get_data", &Matrix::get_data)
.def("multiply", &Matrix::multiply)
.def("__repr__", [](const Matrix& m) {
return "Matrix(" + std::to_string(m.get_rows()) + "x" +
std::to_string(m.get_cols()) + ")";
});
}
Let us walk through this code carefully. The Matrix class stores its data in a one-dimensional vector. For a matrix with r rows and c columns, element at row i and column j is stored at index i times c plus j. This is called row-major order.
The class has two constructors. The first creates a matrix of a given size, initialized with zeros. The second creates a matrix from existing data, checking that the data size matches the dimensions.
The get and set methods provide element access. The get_data method returns a copy of the internal data as a vector, which Pybind11 will convert to a Python list.
The multiply method implements matrix multiplication using the standard algorithm. For each element in the result matrix, we compute the dot product of the corresponding row from the first matrix and column from the second matrix.
In the binding code, we expose both constructors using two py::init declarations. We bind all the methods. We also define repr using a lambda function, which provides a nice string representation when you print a Matrix object in Python.
Compile this module using the appropriate command for your platform. Then you can use it:
import matrix_ops
# Create a 2x3 matrix
m1 = matrix_ops.Matrix(2, 3, [1, 2, 3, 4, 5, 6])
# Create a 3x2 matrix
m2 = matrix_ops.Matrix(3, 2, [7, 8, 9, 10, 11, 12])
# Multiply
result = m1.multiply(m2)
print(f"Result dimensions: {result.get_rows()}x{result.get_cols()}")
print(f"Result data: {result.get_data()}")
This demonstrates a complete, useful C++ class exposed to Python. The matrix multiplication happens entirely in compiled C++ code, which is much faster than pure Python for large matrices.
Comparing Performance
Let us compare the performance of our C++ implementation against pure Python. Create a file named benchmark.py:
import time
import matrix_ops
def python_matrix_multiply(m1_data, m1_rows, m1_cols, m2_data, m2_rows, m2_cols):
if m1_cols != m2_rows:
raise ValueError("Incompatible dimensions")
result = [[0.0 for _ in range(m2_cols)] for _ in range(m1_rows)]
for i in range(m1_rows):
for j in range(m2_cols):
total = 0.0
for k in range(m1_cols):
total += m1_data[i * m1_cols + k] * m2_data[k * m2_cols + j]
result[i][j] = total
return result
# Create test matrices
size = 100
m1_data = [float(i) for i in range(size * size)]
m2_data = [float(i) for i in range(size * size)]
# Benchmark Python
start = time.time()
python_result = python_matrix_multiply(m1_data, size, size, m2_data, size, size)
python_time = time.time() - start
# Benchmark C++
m1 = matrix_ops.Matrix(size, size, m1_data)
m2 = matrix_ops.Matrix(size, size, m2_data)
start = time.time()
cpp_result = m1.multiply(m2)
cpp_time = time.time() - start
print(f"Python time: {python_time:.4f} seconds")
print(f"C++ time: {cpp_time:.4f} seconds")
print(f"Speedup: {python_time / cpp_time:.2f}x")
When you run this benchmark, you will likely see that the C++ version is significantly faster, often by a factor of ten or more. The exact speedup depends on your hardware and compiler optimizations.
This demonstrates why integrating C++ with Python is so valuable. You can write most of your application in Python for productivity and ease of development. Then you identify the performance bottlenecks and implement just those parts in C++.
MEMORY MANAGEMENT AND COMMON PITFALLS
When working with C and C++ from Python, memory management becomes your responsibility in ways it normally is not in pure Python. Understanding the potential pitfalls is crucial to writing robust code.
The Lifetime Problem
Consider this dangerous pattern:
def create_matrix():
data = [1.0, 2.0, 3.0, 4.0]
return matrix_ops.Matrix(2, 2, data)
m = create_matrix()
# m is safe to use here
This is actually fine. The Matrix constructor copies the data from the Python list into its internal storage. The Python list can be garbage collected without affecting the Matrix.
But with ctypes, you could write:
def create_array():
data = [1.0, 2.0, 3.0]
array_type = ctypes.c_double * len(data)
return array_type(*data)
arr = create_array()
# arr is still valid
This is also safe because the ctypes array owns its memory.
The danger comes when you extract pointers or pass data without copying:
# DANGEROUS with some libraries
def get_raw_pointer(arr):
return ctypes.cast(arr, ctypes.POINTER(ctypes.c_double))
If arr goes out of scope after this function returns, the pointer becomes invalid.
Buffer Overruns
C and C++ do not perform bounds checking. If you tell a C function that an array has ten elements, but it actually has five, the function will happily write beyond the end of the array, corrupting memory.
Always ensure that the size you pass to C matches the actual allocated size:
# Correct
data = [1.0, 2.0, 3.0, 4.0, 5.0]
array_type = ctypes.c_double * len(data)
c_array = array_type(*data)
lib.process_array(c_array, len(data)) # Size matches
# WRONG - will corrupt memory
lib.process_array(c_array, 10) # Claiming array is bigger than it is
This is one of the most common sources of crashes when calling C from Python.
Type Mismatches
Passing the wrong type to a C function can cause crashes or silent data corruption. Always specify types explicitly:
# Good
lib.my_function.argtypes = [ctypes.c_int, ctypes.c_double]
lib.my_function(42, 3.14)
# Risky - ctypes will try to convert, but might get it wrong
lib.my_function(42, 3.14) # Without argtypes set
With Pybind11, type checking is automatic and much safer. If you pass the wrong type, you get a clear Python exception rather than undefined behavior.
Memory Leaks
When you allocate memory in C and pass it to Python, you need to ensure it gets freed. With ctypes, if you use malloc in C, you need to call free:
# In C
double* allocate_array(int size) {
return (double*)malloc(size * sizeof(double));
}
void free_array(double* arr) {
free(arr);
}
# In Python
lib.allocate_array.restype = ctypes.POINTER(ctypes.c_double)
lib.free_array.argtypes = [ctypes.POINTER(ctypes.c_double)]
arr = lib.allocate_array(100)
# Use arr...
lib.free_array(arr) # Must remember to free!
This is error-prone. If an exception occurs before free_array is called, you leak memory.
CFFI and Pybind11 handle this better by integrating with Python's garbage collector. When you create objects with ffi.new or return C++ objects from Pybind11 functions, they are automatically freed when no longer referenced.
CHOOSING THE RIGHT TOOL FOR YOUR PROJECT
We have explored three main approaches to calling C and C++ from Python. How do you choose which one to use?
Use ctypes when:
You need to call functions from an existing shared library that you cannot modify. Perhaps you are using a third-party library that provides a C API. The ctypes module is perfect for this because it requires no compilation step. You simply load the library and start calling functions.
You want to avoid dependencies. Since ctypes is part of the standard library, your code will run on any Python installation without requiring additional packages.
Your needs are simple. For basic function calls with primitive types, ctypes is straightforward and gets the job done.
Use CFFI when:
You are working primarily with C code. CFFI's syntax is cleaner than ctypes, and its API mode provides better performance.
You want a good balance between ease of use and performance. CFFI is more modern than ctypes and handles many edge cases better.
You need to work with complex C structures. CFFI makes it easy to define C structs and work with them from Python.
Use Pybind11 when:
You are working with C++ code. Pybind11 is specifically designed for C++ and handles classes, templates, and other C++ features elegantly.
You want the best performance. Pybind11 generates optimized bindings that are often faster than ctypes or CFFI.
You need to expose complex C++ APIs to Python. Pybind11 makes it easy to create Python classes that wrap C++ classes, including inheritance, operator overloading, and template instantiation.
You are building a library that others will use. Pybind11 creates bindings that feel natural to Python users, with proper documentation strings, exception handling, and type checking.
Special Mention: NumPy Integration
If you are working with numerical data, especially arrays, you should consider using NumPy's C API or tools like Pybind11's NumPy integration. NumPy arrays can be passed to C++ code with zero copying, providing excellent performance.
Pybind11 has built-in NumPy support. Include the header:
#include <pybind11/numpy.h>
Then you can write functions that accept NumPy arrays:
#include <pybind11/pybind11.h>
#include <pybind11/numpy.h>
namespace py = pybind11;
double sum_array(py::array_t<double> arr) {
auto buf = arr.request();
double* ptr = static_cast<double*>(buf.ptr);
size_t size = buf.size;
double total = 0.0;
for (size_t i = 0; i < size; i++) {
total += ptr[i];
}
return total;
}
PYBIND11_MODULE(numpy_example, m) {
m.def("sum_array", &sum_array);
}
This function accepts a NumPy array, gets direct access to its underlying data buffer, and sums the elements. No copying occurs. This is extremely efficient for large arrays.
BUILDING A COMPLETE EXAMPLE: LLM TOKEN PROCESSING
Let us conclude with a realistic example that ties everything together. Suppose you are building a Python application that uses a large language model. The model inference is implemented in C++ for speed, but you want to call it from Python.
We will create a simplified token processor that demonstrates the key concepts.
The C++ Token Processor
Create token_processor.cpp:
#include <pybind11/pybind11.h>
#include <pybind11/stl.h>
#include <string>
#include <vector>
#include <unordered_map>
#include <sstream>
namespace py = pybind11;
class TokenProcessor {
private:
std::unordered_map<std::string, int> token_to_id;
std::vector<std::string> id_to_token;
int next_id;
public:
TokenProcessor() : next_id(0) {}
void add_token(const std::string& token) {
if (token_to_id.find(token) == token_to_id.end()) {
token_to_id[token] = next_id;
id_to_token.push_back(token);
next_id++;
}
}
std::vector<int> encode(const std::string& text) {
std::vector<int> result;
std::istringstream stream(text);
std::string word;
while (stream >> word) {
auto it = token_to_id.find(word);
if (it != token_to_id.end()) {
result.push_back(it->second);
} else {
result.push_back(-1);
}
}
return result;
}
std::string decode(const std::vector<int>& token_ids) {
std::string result;
for (size_t i = 0; i < token_ids.size(); i++) {
if (i > 0) {
result += " ";
}
int id = token_ids[i];
if (id >= 0 && id < static_cast<int>(id_to_token.size())) {
result += id_to_token[id];
} else {
result += "[UNK]";
}
}
return result;
}
int vocab_size() const {
return next_id;
}
};
PYBIND11_MODULE(token_processor, m) {
m.doc() = "Fast token processing for LLM applications";
py::class_<TokenProcessor>(m, "TokenProcessor")
.def(py::init<>())
.def("add_token", &TokenProcessor::add_token,
"Add a token to the vocabulary")
.def("encode", &TokenProcessor::encode,
"Encode text into token IDs")
.def("decode", &TokenProcessor::decode,
"Decode token IDs back into text")
.def("vocab_size", &TokenProcessor::vocab_size,
"Get the current vocabulary size");
}
This class maintains a vocabulary of tokens. It can encode text into token IDs and decode token IDs back into text. This is a simplified version of what real tokenizers do.
The implementation uses a hash map for fast token-to-ID lookup and a vector for ID-to-token lookup. The encode method splits the input text by whitespace and looks up each word. The decode method reconstructs text from token IDs.
Using the Token Processor in Python
After compiling the module, you can use it like this:
import token_processor
# Create a processor
processor = token_processor.TokenProcessor()
# Build vocabulary
vocabulary = ["hello", "world", "python", "is", "awesome"]
for token in vocabulary:
processor.add_token(token)
print(f"Vocabulary size: {processor.vocab_size()}")
# Encode some text
text = "hello world python is awesome"
token_ids = processor.encode(text)
print(f"Encoded: {token_ids}")
# Decode back
decoded = processor.decode(token_ids)
print(f"Decoded: {decoded}")
# Try with unknown tokens
text_with_unknown = "hello universe python is great"
token_ids = processor.encode(text_with_unknown)
print(f"With unknowns: {token_ids}")
decoded = processor.decode(token_ids)
print(f"Decoded with unknowns: {decoded}")
This demonstrates a complete workflow. We create the processor, build a vocabulary, encode text into IDs, and decode IDs back into text. Unknown tokens are handled gracefully.
The key point is that all the heavy lifting happens in C++. The hash map lookups, string operations, and vector manipulations are all compiled code running at native speed. Python just orchestrates the high-level logic.
Extending with Batch Processing
Real LLM applications often process batches of text for efficiency. Let us extend our processor to handle batches:
Add to token_processor.cpp, inside the TokenProcessor class:
std::vector<std::vector<int>> encode_batch(const std::vector<std::string>& texts) {
std::vector<std::vector<int>> results;
results.reserve(texts.size());
for (const auto& text : texts) {
results.push_back(encode(text));
}
return results;
}
Add to the binding code:
.def("encode_batch", &TokenProcessor::encode_batch,
"Encode multiple texts into token IDs")
Now you can process multiple texts at once:
texts = [
"hello world",
"python is awesome",
"hello python"
]
batch_results = processor.encode_batch(texts)
for i, token_ids in enumerate(batch_results):
print(f"Text {i}: {token_ids}")
This batch processing happens entirely in C++, avoiding the overhead of repeatedly crossing the Python-C++ boundary.
FINAL THOUGHTS AND BEST PRACTICES
You now have a solid foundation for integrating C and C++ code with Python. Let us review some best practices to keep in mind.
Always start simple. Begin with a minimal example to verify that your build process works. Then gradually add complexity. This makes debugging much easier.
Document your type signatures carefully. Whether using ctypes argtypes, CFFI cdef declarations, or Pybind11 bindings, clear type information prevents bugs and makes your code maintainable.
Test thoroughly at the boundaries. The interface between Python and C++ is where bugs often hide. Write tests that verify correct behavior with various input types, edge cases, and error conditions.
Handle errors properly. C++ exceptions can be caught and converted to Python exceptions by Pybind11 automatically. With ctypes and CFFI, you need to check return values and handle errors explicitly.
Profile before optimizing. Do not assume that moving code to C++ will make it faster. Profile your Python code first to identify actual bottlenecks. Sometimes algorithmic improvements in Python are more effective than rewriting in C++.
Consider maintenance costs. C++ code is harder to write and debug than Python. Only use C++ for parts of your application where the performance benefit justifies the added complexity.
Use modern C++ features. If you are writing new C++ code, use C++11 or later. Modern C++ is safer and more expressive than older versions. Pybind11 requires C++11 anyway.
Leverage existing libraries. Before writing your own C++ code, check if a library already exists. For numerical computing, consider NumPy and SciPy. For machine learning, look at existing frameworks. Standing on the shoulders of giants is always preferable to reinventing wheels.
Keep the interface simple. The simpler your C++ API, the easier it is to bind to Python. Avoid complex template metaprogramming at the interface boundary. Use concrete types where possible.
Version your bindings carefully. When you update your C++ code, ensure backward compatibility or clearly version your Python module. Breaking changes in compiled extensions are harder to manage than pure Python code.
CONCLUSION
The ability to call C and C++ from Python is a superpower for developers. It allows you to combine Python's productivity with the performance of compiled languages. Whether you are implementing fast numerical algorithms, integrating with existing C libraries, or building high-performance components for machine learning applications, these techniques are essential tools in your toolkit.
We have covered three main approaches: ctypes for simple C function calls, CFFI for a more modern C interface, and Pybind11 for elegant C++ integration. Each has its place, and understanding all three allows you to choose the right tool for each situation.
The examples we have worked through, from simple addition functions to matrix multiplication to token processing, demonstrate the patterns you will use in real projects. The key is understanding the flow of data between Python and compiled code, managing memory correctly, and choosing appropriate abstractions.
As you build your own applications, remember that integration is a means to an end. The goal is not to use C++ for its own sake, but to create better software. Use Python where it excels: rapid development, clear code, rich libraries. Use C++ where it excels: computational performance, low-level control, integration with existing systems.
With the knowledge from this tutorial, you are equipped to build sophisticated applications that leverage the strengths of both languages. Whether you are accelerating machine learning inference, processing large datasets, or building high-performance APIs, you now have the tools and understanding to succeed.