Language Integration: C & Rust
Reviewed & published by Brayan K
Master Python + C/Rust integration using ctypes, cffi, Cython, and PyO3 for maximum performance while maintaining Python's productivity
Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn
| Method | Difficulty | Best For | Speed Gain |
|---|---|---|---|
| ctypes | Easy | Quick prototyping, calling existing C libs | 10-50x |
| CFFI | Medium | Cleaner C interface, better error handling | 10-50x |
| Cython | Medium | Gradual optimization of Python code | 10-100x |
| PyO3 (Rust) | Hard | Memory-safe, blazing fast extensions | 50-1000x |
- Why and when to integrate Python with other languages
- CPython's architecture and extension mechanism
- Using ctypes to call C libraries without compilation
- cffi for more ergonomic C interfacing
- Writing native Python extensions in C/C++
- Cython for Python-like syntax with C performance
- Rust integration with PyO3 for memory safety
- Embedding Python in C/Rust applications
- Data marshalling and zero-copy strategies
- Performance optimization and profiling
- Safety considerations and testing
- Building and distributing cross-platform extensions
Why Integrate Python with Other Languages?
Python is excellent for productivity, but sometimes you need more. Language integration lets you keep Python's ease of use while accessing native performance and capabilities.
Common Reasons
- Performance - Move CPU-intensive loops to C/Rust (10-100x speedup)
- Existing libraries - Use mature C/C++/Rust libraries without rewriting
- System-level access - Hardware drivers, OS syscalls, low-level protocols
- Memory safety - Rust prevents segfaults and memory leaks
- Incremental optimization - Profile first, optimize only bottlenecks
The 95/5 Rule
Keep 95% of your code in Python for maintainability. Move the critical 5% (hot loops, numeric kernels) to native code. This is exactly how NumPy, pandas, and PyTorch work.
CPython's Architecture
Understanding CPython's internals helps you write better extensions.
Key Facts
- CPython itself is written in C
- All Python objects are C structs (PyObject, PyListObject, etc.)
- The GIL (Global Interpreter Lock) serializes Python bytecode execution
- C extensions can release the GIL for CPU-bound work
- The Python/C API lets you create new types and call Python from C
How Extensions Work
When you import myextension, Python:
- Looks for a shared library (.so, .pyd, .dll)
- Loads it dynamically
- Calls PyInit_myextension()
- Registers functions and types with the Python runtime
ctypes: Call C Without Compilation
ctypes is Python's built-in FFI (Foreign Function Interface) for calling C libraries.
Workflow
- Compile C code as shared library (.so/.dll/.dylib)
- Load with ctypes.CDLL()
- Declare function argument and return types
- Call functions like normal Python
✓ Advantages
- • No Python-specific build step
- • Works with existing binaries
- • Part of Python standard library
- • Cross-platform
✗ Limitations
- • Manual type declarations
- • Easy to make mistakes (crashes)
- • Limited struct/pointer support
- • No automatic error handling
cffi: More Ergonomic C Interface
cffi improves on ctypes with C-style declarations and better error handling.
Key Features
- C-style API definitions in Python strings
- Better struct, pointer, and enum support
- Two modes: ABI (existing binaries) and API (compile from source)
- Automatic type conversions
- Used by popular libraries like cryptography and PyPy
When to Use
Choose cffi when you need to wrap complex C APIs with many structs and pointers, or when you want more safety than raw ctypes provides.
Native Python Extensions in C/C++
The Python/C API provides maximum control and performance for extension modules.
Core Components
- PyObject* - Base type for all Python objects
- PyArg_ParseTuple - Parse function arguments
- Py_BuildValue - Construct return values
- PyMethodDef - Method table for module functions
- PyModuleDef - Module definition
- PyInit_* - Module initialization function
Use Cases
- NumPy-style array operations
- Custom Python types in C
- Maximum performance numeric code
- Direct CPython runtime manipulation
⚠️ Warning
C extensions are powerful but error-prone. Manual memory management, reference counting, and GIL handling make them challenging. Consider Cython or Rust instead for new projects.
Cython: Python Syntax, C Speed
Cython lets you write Python-like code with optional type annotations that compile to C.
- Familiar syntax - Looks like Python with type hints
- Gradual typing - Add types where they matter for speed
- C integration - Call C functions directly
- NumPy support - Efficient typed array operations
- Compiler directives - Disable bounds checking, wraparound
Performance Tips
• Use cdef for C-level variables and functions
• Use typed memory views for arrays: double[:]
• Release the GIL for CPU-intensive work: with nogil:
Perfect For
Scientific computing, numeric algorithms, and data processing. If your bottleneck is a tight numeric loop, Cython is often the fastest path to optimization.
Rust Integration with PyO3
Rust offers C-like performance with memory safety guarantees. PyO3 makes Rust-Python integration seamless.
Why Rust + Python?
- Memory safety - No segfaults, no data races
- Performance - Comparable to C/C++
- Modern tooling - Cargo, rustfmt, clippy
- Rich ecosystem - Excellent libraries for crypto, networking, parsing
- Fearless concurrency - Safe parallel processing
Functions
- • #[pyfunction] macro
- • Automatic type conversion
- • Python exceptions
- • Default arguments
Classes
- • #[pyclass] for types
- • #[pymethods] for methods
- • Properties and attributes
- • Magic methods (__repr__, etc.)
Embedding Python
Sometimes you want to embed Python as a scripting language inside a C/Rust application.
- Game engines - Let users write Python mods and scripts
- Plugin systems - Extensible applications with Python plugins
- Configuration - Python as a config language
- Trading systems - Rust/C core with Python strategies
- Scientific tools - Fast core with Python analysis
Basic Workflow
- Initialize Python runtime: Py_Initialize()
- Import modules and run code
- Call Python functions from C/Rust
- Pass data between languages
- Cleanup: Py_Finalize()
Data Marshalling Strategies
Efficient data transfer between languages is critical for performance.
Zero-Copy Techniques
- NumPy arrays - Pass raw pointers to contiguous memory
- Buffer protocol - Python's standard for binary data
- Memory views - Share memory without copying
- Typed memory views in Cython - double[:]
Best Practices
• Batch operations - Call once with 1M elements, not 1M times with 1 element
• Use simple types at boundaries - Primitives, strings, byte arrays
• Avoid repeated conversions - Convert once, reuse
• Check array layout - Ensure C-contiguous for efficient access
⚠️ Common Mistake
Converting large NumPy arrays to Python lists loses all performance benefits. Always pass array pointers directly to C/Rust code.
🧑🏫 Worked example: see the bytes for yourself
You can't compile a C extension in a browser tab, but you can do the part that actually confuses people: looking at the raw bytes that cross the boundary. Everything below is pure standard library, so it runs anywhere. struct turns Python values into the exact byte layout a C struct expects, array gives you one contiguous block instead of scattered Python objects, and memoryview is the zero-copy window — the thing "pass a pointer" really means in Python.
# WORKED EXAMPLE: what actually crosses the boundary when Python calls C or Rust.
# C has no idea what a Python list is. It wants a flat block of bytes at an address.
# struct, array and memoryview let you see - and control - that block.
import struct
from array import array
# --- 1. Single values: struct.pack turns Python values into raw bytes --------
# "<" = little-endian byte order (what x86 and ARM use)
# "i" = 32-bit signed int, "f" = 32-bit float
packed = struct.pack("<if", 42, 1.5)
print("bytes going to C:", packed.hex(" ")) # 8 bytes: 4 for the int, 4 for the float
print("length in bytes:", len(packed))
# struct.unpack reverses it - this is how you read a struct C hands back
print("back in Python:", struct.unpack("<if", packed))
# --- 2. Many values: array gives a contiguous C block, a Python list does not -
numbers = array("d", [1.0, 2.0, 3.0, 4.0]) # "d" = C double, 8 bytes each
print("bytes per item:", numbers.itemsize)
print("total bytes:", len(numbers) * numbers.itemsize)
# --- 3. Zero-copy: memoryview points at the SAME bytes, it does not duplicate -
view = memoryview(numbers)
print("view format:", view.format, "| items:", len(view))
view[0] = 99.0 # write through the view...
print("original array is now:", numbers.tolist()) # ...and the array changed too
# --- 4. The expensive mistake this lesson warns about ------------------------
copied = list(numbers) # builds 4 separate Python float objects, one allocation each
print("as a Python list:", copied)
print("fine for 4 items; ruinous for 4 million - that is why you pass the buffer")
# ✅ Output:
# bytes going to C: 2a 00 00 00 00 00 c0 3f
# length in bytes: 8
# back in Python: (42, 1.5)
# bytes per item: 8
# total bytes: 32
# view format: d | items: 4
# original array is now: [99.0, 2.0, 3.0, 4.0]
# as a Python list: [99.0, 2.0, 3.0, 4.0]
# fine for 4 items; ruinous for 4 million - that is why you pass the buffer
# ✅ Expected output:
# bytes going to C: 2a 00 00 00 00 00 c0 3f
# length in bytes: 8
# back in Python: (42, 1.5)
# bytes per item: 8
# total bytes: 32
# view format: d | items: 4
# original array is now: [99.0, 2.0, 3.0, 4.0]
# as a Python list: [99.0, 2.0, 3.0, 4.0]
# fine for 4 items; ruinous for 4 million - that is why you pass the buffer🎯 Your turn: marshal a sensor reading
Same tools, your hands. A C function wants a Reading struct — a 32-bit int followed by a 32-bit float — and then a batch of doubles it can write back into. Fill in the three ___ blanks and run it.
# 🎯 YOUR TURN - marshal a record for a C function
#
# The C side declares:
# struct Reading {
# int sensor_id; /* 32-bit signed integer */
# float temperature; /* 32-bit float */
# };
import struct
from array import array
# 1) Build the format string: "<" for little-endian, then the code for a
# 32-bit int, then the code for a 32-bit float.
FORMAT = ___ # 👉 replace ___ with that 3-character string, in quotes
record = struct.pack(FORMAT, 7, 21.5)
print("packed:", record.hex(" "))
print("size:", struct.calcsize(FORMAT), "bytes")
print("unpacked:", struct.unpack(FORMAT, record))
# 2) Now a whole batch of readings as C doubles.
readings = array(___, [21.5, 22.0, 22.5, 23.0]) # 👉 replace ___ with the type code for a C double, in quotes
print("batch bytes:", len(readings) * readings.itemsize)
# 3) Hand C a zero-copy window onto that batch instead of a fresh copy.
buffer = ___(readings) # 👉 replace ___ with the built-in that shares memory
buffer[3] = 30.0 # pretend the C function wrote a corrected value here
print("after C wrote to the buffer:", readings.tolist())
# ✅ Expected output once the blanks are filled in:
# packed: 07 00 00 00 00 00 ac 41
# size: 8 bytes
# unpacked: (7, 21.5)
# batch bytes: 32
# after C wrote to the buffer: [21.5, 22.0, 22.5, 30.0]
#
# If the last line still ends in 23.0, you copied instead of sharing - step 3 is wrong.When to Choose Which Integration Path
Decision Matrix
→ Cython or NumPy optimization
Calling existing C library?
→ ctypes (simple) or cffi (complex APIs)
→ Rust + PyO3 (modern) or C++ with pybind11 (legacy)
Need scripting in native app?
→ Embed Python with CPython API or PyO3
→ Direct C extension (but consider alternatives first)
Practical Workflow
The best approach is incremental optimization based on profiling.
Recommended Process
- Write everything in pure Python Get it working first. Premature optimization wastes time.
- Profile to find bottlenecks Use cProfile, line_profiler, or py-spy to identify hot spots.
- Optimize in Python first Use NumPy, better algorithms, caching. Native code may not be needed.
- Move critical functions to native code Only the 5-10% that's truly slow. Keep the rest in Python.
- Wrap behind Python API Callers shouldn't know it's implemented in C/Rust.
- Add comprehensive tests Test the boundary thoroughly. Native code bugs are harder to debug.
Performance & Safety
Critical Considerations
- Memory management - Track lifetimes, avoid use-after-free
- Reference counting - Properly incref/decref in C extensions
- GIL handling - Release for CPU work, hold for Python API calls
- Error propagation - Convert native errors to Python exceptions
- Type safety - Validate inputs at language boundaries
- Thread safety - Native code may run without GIL protection
Unit Tests
Test native code in isolation with C/Rust test frameworks
Integration Tests
Test Python API with pytest, validate behavior and errors
Fuzz Testing
Generate random inputs to find crashes and edge cases
Building & Distribution
Professional packages need cross-platform wheels that users can install without compilers.
For Rust (PyO3)
maturin - Handles building, packaging, and uploading to PyPI
For C/C++
setuptools with Extension or scikit-build for CMake
CI/CD for Extensions
Use GitHub Actions with cibuildwheel or maturin to build wheels for:
- • Linux: manylinux (x86_64, aarch64)
- • macOS: Intel and Apple Silicon
- • Windows: 32-bit and 64-bit
Common Pitfalls
✗ Don't Do This
- Optimize before profiling - you'll optimize the wrong thing
- Convert large arrays to Python lists - destroys performance
- Call native functions in tight Python loops - batch instead
- Ignore error handling at boundaries - leads to crashes
- Store pointers to Python memory in native code - lifecycle issues
- Use shell=True in subprocess - security risk
- Build giant monolithic native modules - hard to maintain
✓ Best Practices
- Profile first, optimize hot paths only
- Use NumPy array pointers for zero-copy data passing
- Keep native modules focused and small
- Test boundaries thoroughly with edge cases
- Validate inputs before crossing to native code
- Document memory ownership clearly
- Provide fallback Python implementations
# Python Language Integration Examples
# ============================================
# 1. CALLING C WITH CTYPES (NO COMPILATION)
# ============================================
import ctypes
import os
from pathlib import Path
# Example 1: Basic ctypes usage
def load_c_library_example():
"""Load and call functions from a shared C library"""
# On Linux/macOS: libmath.so, on Windows: math.dll
lib_name = "libmath.so" if os.name != "nt" else "math.dll"
try:
# Load the library
lib = ctypes.CDLL(lib_name)
# Define function signatures
lib.add.argtypes = (ctypes.c_double, ctypes.c_double)
lib.add.restype = ctypes.c_double
lib.multiply.argtypes = (ctypes.c_int, ctypes.c_int)
lib.multiply.restype = ctypes.c_int
# Call the functions
result1 = lib.add(3.5, 2.5)
result2 = lib.multiply(7, 6)
print(f"C add(3.5, 2.5) = {result1}")
print(f"C multiply(7, 6) = {result2}")
except OSError as e:
print(f"Could not load library: {e}")
# Example 2: Passing arrays to C
def array_to_c_example():
"""Pass NumPy arrays to C functions"""
import numpy as np
# Simulate loading a C library
# In reality: lib = ctypes.CDLL("./libarray.so")
# Create NumPy array
arr = np.array([1.0, 2.0, 3.0, 4.0, 5.0], dtype=np.float64)
# Get pointer to array data
arr_ptr = arr.ctypes.data_as(ctypes.POINTER(ctypes.c_double))
size = arr.size
print(f"Array: {arr}")
print(f"Array pointer: {arr_ptr}")
print(f"Array size: {size}")
# In C, function signature would be:
# double sum_array(double* arr, int size);
# Example 3: Working with structures
class Point(ctypes.Structure):
"""C struct representation"""
_fields_ = [
("x", ctypes.c_double),
("y", ctypes.c_double)
]
def struct_example():
"""Work with C structures"""
p1 = Point(3.0, 4.0)
p2 = Point(1.0, 2.0)
print(f"Point 1: ({p1.x}, {p1.y})")
print(f"Point 2: ({p2.x}, {p2.y})")
# Pass to C function that computes distance
# In C: double distance(Point* p1, Point* p2);
# ============================================
# 2. CFFI - MORE ERGONOMIC C INTERFACE
# ============================================
def cffi_example():
"""Using cffi for cleaner C interface"""
try:
from cffi import FFI
ffi = FFI()
# Define C interface
ffi.cdef("""
double add(double a, double b);
double multiply(double a, double b);
typedef struct {
double x;
double y;
} Point;
double distance(Point* p1, Point* p2);
""")
# Load library
# C = ffi.dlopen("./libmath.so")
# In production:
# result = C.add(3.5, 2.5)
# print(f"Result: {result}")
print("cffi example ready (requires compiled C library)")
except ImportError:
print("cffi not installed. Install with: pip install cffi")
# ============================================
# 3. PYTHON/C API EXTENSION EXAMPLE
# ============================================
# This would be in a .c file, shown here for reference
c_extension_code = '''
// Example C extension module
#include <Python.h>
static PyObject* fast_sum(PyObject* self, PyObject* args) {
PyObject* list_obj;
if (!PyArg_ParseTuple(args, "O!", &PyList_Type, &list_obj)) {
return NULL;
}
Py_ssize_t size = PyList_Size(list_obj);
double sum = 0.0;
for (Py_ssize_t i = 0; i < size; i++) {
PyObject* item = PyList_GetItem(list_obj, i);
double value = PyFloat_AsDouble(item);
if (PyErr_Occurred()) {
return NULL;
}
sum += value;
}
return PyFloat_FromDouble(sum);
}
static PyMethodDef Methods[] = {
{"fast_sum", fast_sum, METH_VARARGS, "Sum a list of numbers"},
{NULL, NULL, 0, NULL}
};
static struct PyModuleDef moduledef = {
PyModuleDef_HEAD_INIT,
"cextension",
"Example C extension",
-1,
Methods
};
PyMODINIT_FUNC PyInit_cextension(void) {
return PyModule_Create(&moduledef);
}
'''
print("C Extension example (reference only):")
print(c_extension_code)
# ============================================
# 4. CYTHON EXAMPLE
# ============================================
# This would be in a .pyx file
cython_code = '''
# cython: boundscheck=False, wraparound=False
# File: fastmath.pyx
cimport cython
from libc.math cimport sqrt
def fast_distance(double[:] x1, double[:] y1,
double[:] x2, double[:] y2):
"""Compute distances between point arrays"""
cdef Py_ssize_t i, n = x1.shape[0]
cdef double[:] result = np.empty(n, dtype=np.float64)
cdef double dx, dy
for i in range(n):
dx = x2[i] - x1[i]
dy = y2[i] - y1[i]
result[i] = sqrt(dx * dx + dy * dy)
return np.asarray(result)
@cython.cdivision(True)
def fast_mean(double[:] arr):
"""Fast mean calculation"""
cdef Py_ssize_t i, n = arr.shape[0]
cdef double total = 0.0
for i in range(n):
total += arr[i]
return total / n
cdef class FastVector:
"""Fast vector operations"""
cdef double x, y, z
def __init__(self, double x, double y, double z):
self.x = x
self.y = y
self.z = z
cpdef double magnitude(self):
return sqrt(self.x*self.x + self.y*self.y + self.z*self.z)
cpdef FastVector add(self, FastVector other):
return FastVector(
self.x + other.x,
self.y + other.y,
self.z + other.z
)
'''
print("\nCython example (would be compiled):")
print(cython_code)
# ============================================
# 5. RUST + PyO3 EXAMPLE
# ============================================
rust_code = '''
// Rust code for Python integration using PyO3
// File: src/lib.rs
use pyo3::prelude::*;
use pyo3::types::PyList;
#[pyfunction]
fn add(a: f64, b: f64) -> f64 {
a + b
}
#[pyfunction]
fn fast_sum(numbers: Vec<f64>) -> f64 {
numbers.iter().sum()
}
#[pyfunction]
fn multiply_array(mut arr: Vec<f64>, factor: f64) -> Vec<f64> {
for x in arr.iter_mut() {
*x *= factor;
}
arr
}
#[pyclass]
struct Point {
#[pyo3(get, set)]
x: f64,
#[pyo3(get, set)]
y: f64,
}
#[pymethods]
impl Point {
#[new]
fn new(x: f64, y: f64) -> Self {
Point { x, y }
}
fn distance(&self, other: &Point) -> f64 {
let dx = self.x - other.x;
let dy = self.y - other.y;
(dx * dx + dy * dy).sqrt()
}
fn __repr__(&self) -> String {
format!("Point({}, {})", self.x, self.y)
}
}
#[pymodule]
fn rustmath(_py: Python<'_>, m: &PyModule) -> PyResult<()> {
m.add_function(wrap_pyfunction!(add, m)?)?;
m.add_function(wrap_pyfunction!(fast_sum, m)?)?;
m.add_function(wrap_pyfunction!(multiply_array, m)?)?;
m.add_class::<Point>()?;
Ok(())
}
// Cargo.toml would include:
// [dependencies]
// pyo3 = { version = "0.22", features = ["extension-module"] }
//
// [lib]
// crate-type = ["cdylib"]
'''
print("\nRust + PyO3 example:")
print(rust_code)
# ============================================
# 6. EMBEDDING PYTHON IN C/RUST
# ============================================
embedding_c_code = '''
// Embedding Python in C
#include <Python.h>
int main() {
Py_Initialize();
// Run simple Python code
PyRun_SimpleString("print('Hello from embedded Python')");
// Import a module and call a function
PyObject* pModule = PyImport_ImportModule("math");
if (pModule) {
PyObject* pFunc = PyObject_GetAttrString(pModule, "sqrt");
if (pFunc && PyCallable_Check(pFunc)) {
PyObject* pArgs = Py_BuildValue("(d)", 25.0);
PyObject* pValue = PyObject_CallObject(pFunc, pArgs);
if (pValue) {
double result = PyFloat_AsDouble(pValue);
printf("sqrt(25) = %f\\n", result);
Py_DECREF(pValue);
}
Py_DECREF(pArgs);
Py_DECREF(pFunc);
}
Py_DECREF(pModule);
}
Py_Finalize();
return 0;
}
'''
embedding_rust_code = '''
// Embedding Python in Rust with PyO3
use pyo3::prelude::*;
use pyo3::types::IntoPyDict;
fn main() -> PyResult<()> {
Python::with_gil(|py| {
// Execute Python code
py.run("print('Hello from Rust!')", None, None)?;
// Import and use Python modules
let math = py.import("math")?;
let result: f64 = math.getattr("sqrt")?.call1((25.0,))?.extract()?;
println!("sqrt(25) = {}", result);
// Create Python objects from Rust
let locals = [("value", 42)].into_py_dict(py);
py.run(
"print(f'The value is {value}')",
None,
Some(locals),
)?;
Ok(())
})
}
'''
print("\nEmbedding Python in C:")
print(embedding_c_code)
print("\nEmbedding Python in Rust:")
print(embedding_rust_code)
# ============================================
# 7. PERFORMANCE COMPARISON EXAMPLE
# ============================================
import time
import numpy as np
def pure_python_sum(numbers):
"""Pure Python summation"""
total = 0.0
for num in numbers:
total += num
return total
def numpy_sum(numbers):
"""NumPy summation (C-optimized)"""
return np.sum(numbers)
def benchmark_implementations():
"""Compare performance of different implementations"""
size = 1_000_000
data = list(range(size))
np_data = np.array(data, dtype=np.float64)
# Pure Python
start = time.perf_counter()
result1 = pure_python_sum(data)
python_time = time.perf_counter() - start
# NumPy (C-optimized)
start = time.perf_counter()
result2 = numpy_sum(np_data)
numpy_time = time.perf_counter() - start
print(f"\nPerformance Comparison ({size:,} elements):")
print(f"Pure Python: {python_time:.4f}s")
print(f"NumPy (C): {numpy_time:.4f}s")
print(f"Speedup: {python_time / numpy_time:.1f}x")
print(f"Results match: {abs(result1 - result2) < 1e-9}")
# ============================================
# 8. PRACTICAL WORKFLOW EXAMPLE
# ============================================
class HybridProcessor:
"""Example of Python-first, selectively optimized approach"""
def __init__(self, use_native: bool = False):
self.use_native = use_native
if use_native:
try:
# Try to import compiled native module
# import fastnative
self.native_available = False # Would be True if import succeeded
except ImportError:
self.native_available = False
else:
self.native_available = False
def process_data(self, data: np.ndarray) -> np.ndarray:
"""Process data with optional native acceleration"""
if self.native_available:
# Use fast native implementation
# return fastnative.process(data)
pass
# Fallback to pure Python/NumPy
return self._python_process(data)
def _python_process(self, data: np.ndarray) -> np.ndarray:
"""Pure Python implementation"""
# Data normalization
mean = np.mean(data)
std = np.std(data)
normalized = (data - mean) / (std + 1e-10)
# Apply transformation
transformed = np.tanh(normalized)
return transformed
# ============================================
# 9. DATA MARSHALLING STRATEGIES
# ============================================
def efficient_data_passing():
"""Demonstrate efficient data passing between Python and C"""
import numpy as np
# Create large array
data = np.random.rand(1_000_000)
# INEFFICIENT: Converting to Python list
# python_list = data.tolist() # Creates copy
# pass_to_c(python_list) # Another conversion
# EFFICIENT: Pass NumPy array pointer directly
# Via ctypes:
data_ptr = data.ctypes.data_as(ctypes.POINTER(ctypes.c_double))
size = data.size
print("Efficient data passing:")
print(f"Array size: {size:,}")
print(f"Array dtype: {data.dtype}")
print(f"Contiguous: {data.flags.c_contiguous}")
print(f"Pointer: {data_ptr}")
# In C, this would be:
# void process(double* data, size_t size) { ... }
# ============================================
# 10. INTEGRATION DECISION HELPER
# ============================================
def suggest_integration_approach(
bottleneck_type: str,
complexity: str,
team_expertise: str
):
"""Help decide which integration approach to use"""
recommendations = {
"numeric_simple": {
"python": "NumPy or Cython",
"c": "Cython (Python-like syntax)",
"rust": "Not necessary for simple numeric code"
},
"numeric_complex": {
"python": "Cython with custom types",
"c": "C extension or Cython",
"rust": "Rust + PyO3 for safety"
},
"systems": {
"python": "Use subprocess or ctypes",
"c": "C extension with proper error handling",
"rust": "Rust + PyO3 (best safety/performance)"
},
"existing_lib": {
"python": "ctypes or cffi",
"c": "Direct C extension wrapper",
"rust": "Use bindgen + PyO3"
}
}
print(f"\nIntegration Recommendations:")
print(f"Bottleneck: {bottleneck_type}")
print(f"Complexity: {complexity}")
print(f"Team expertise: {team_expertise}")
print("\nSuggested approaches:")
if bottleneck_type in recommendations:
for lang, approach in recommendations[bottleneck_type].items():
print(f" {lang.upper()}: {approach}")
# Run examples
print("Python Language Integration Examples\n")
print("=" * 60)
load_c_library_example()
array_to_c_example()
struct_example()
cffi_example()
benchmark_implementations()
efficient_data_passing()
suggest_integration_approach("numeric_complex", "high", "python")
print("\nExamples completed!")🎯 Mini Challenge: Decode a Native Frame
Native code rarely hands you a tidy Python string. It hands you a frame: a small binary header saying how many bytes follow, then the bytes. Write both halves — the encoder that builds a frame and the decoder that reads one back. This challenge also lands the boundary bug that catches almost everyone the first time: character count and byte count are not the same number.
The starter is an outline only. Use the print formats given and your output will match exactly.
# 🎯 MINI-CHALLENGE: a length-prefixed binary frame
#
# The wire format your Rust extension uses:
# [ 2-byte unsigned little-endian length ][ that many UTF-8 bytes ]
#
# struct format code you need: "<H" ("H" = 16-bit unsigned int, 2 bytes)
import struct
# 1. Write encode(text) -> bytes
# - encode text to UTF-8 bytes
# - pack the NUMBER OF BYTES with "<H" and put it in front
# - return header + data
#
# 2. Write decode(frame) -> (text, byte_count)
# - unpack the first 2 bytes to get the length (unpack always returns a tuple)
# - slice out exactly that many bytes after the header and decode them as UTF-8
# - return the text and the byte count
#
# 3. Round-trip the word "héllo" and print these five lines:
# f"frame: {frame.hex(' ')}"
# f"length header says: {n} bytes"
# f"decoded: {text!r}"
# f"total frame size: {len(frame)} bytes"
# f"characters: {len(text)} bytes: {n}"
# your code here
# ✅ Expected output:
# frame: 06 00 68 c3 a9 6c 6c 6f
# length header says: 6 bytes
# decoded: 'héllo'
# total frame size: 8 bytes
# characters: 5 bytes: 6
#
# That last line is the whole point: "héllo" is 5 characters but 6 bytes, because
# é needs two. Size a native buffer with len(text) and you will corrupt the message.Key Takeaways
- Python + C/Rust integration gives you productivity AND performance
- Keep 95% in Python, optimize the critical 5% in native code
- ctypes is easiest for simple C libraries, cffi for complex ones
- Cython offers Python-like syntax with C performance for numeric code
- Rust + PyO3 provides memory safety and modern tooling
- Always profile before optimizing - intuition is often wrong
- Zero-copy data passing via NumPy pointers is critical for performance
- Test language boundaries thoroughly - bugs are harder to debug
- Build cross-platform wheels for easy pip install experience
- This is how NumPy, pandas, PyTorch, and most fast libraries work
📋 Quick Reference — Language Integration
| Tool | What it does |
|---|---|
| ctypes.CDLL('lib.so') | Load and call a C shared library |
| ctypes.c_int / ctypes.c_double | C type mappings for Python |
| cffi | Modern C integration via inline C headers |
| PyO3 (Rust) | Write Python extensions in Rust |
| numpy array as C pointer | Zero-copy data sharing with C/Rust |
🎉 Great work! You've completed this lesson.
You can now call C and Rust from Python — the same technique used in NumPy, pandas, PyTorch, and every high-performance Python library.
Practice quiz
Which integration tool is built into Python's standard library and needs no compilation step to call an existing C library?
- Cython
- PyO3
- ctypes
- maturin
Answer: ctypes. ctypes is Python's built-in FFI; it loads a shared library with CDLL() and needs no Python-specific build step.
What is the '95/5 Rule' described in the lesson?
- Keep 95% of code in Python and move the critical 5% to native code
- Keep 95% in native code, 5% in Python
- Optimize 95% of functions before profiling
- Use 95 native modules per project
Answer: Keep 95% of code in Python and move the critical 5% to native code. Keep 95% in Python for maintainability and move only the critical 5% (hot loops) to native code.
What is the GIL in CPython?
- A garbage collector
- A C compiler
- A memory allocator
- The Global Interpreter Lock that serializes Python bytecode execution
Answer: The Global Interpreter Lock that serializes Python bytecode execution. The GIL (Global Interpreter Lock) serializes Python bytecode; C extensions can release it for CPU-bound work.
Which tool offers Python-like syntax with optional type annotations that compile to C?
- ctypes
- Cython
- cffi
- PyO3
Answer: Cython. Cython lets you write Python-like code with optional types that compile to C for speed.
Which language binding emphasizes memory safety with no segfaults or data races?
- Rust + PyO3
- C extension via the Python/C API
- ctypes
- cffi
Answer: Rust + PyO3. Rust offers memory safety (no segfaults, no data races) and PyO3 makes Rust-Python integration seamless.
Which build tool handles building, packaging, and uploading PyO3 (Rust) extensions to PyPI?
- setuptools
- cythonize
- maturin
- cibuildwheel
Answer: maturin. maturin builds, packages, and publishes Rust/PyO3 extension wheels.
What is the recommended FIRST step in the practical optimization workflow?
- Rewrite everything in C immediately
- Write everything in pure Python and get it working first
- Add PyO3 bindings
- Disable the GIL
Answer: Write everything in pure Python and get it working first. Write it in pure Python first; premature optimization wastes time. Profile before optimizing.
What is the most efficient way to pass a large NumPy array to C/Rust code?
- Convert it to a Python list first
- Pickle the array
- Send it one element at a time
- Pass the raw array pointer directly (zero-copy)
Answer: Pass the raw array pointer directly (zero-copy). Zero-copy: pass the array's data pointer directly. Converting to a Python list destroys performance.
What does ctypes.CDLL('lib.so') do?
- Compiles a C source file
- Loads a C shared library so you can call its functions
- Creates a new Python class
- Releases the GIL
Answer: Loads a C shared library so you can call its functions. CDLL loads and gives access to a C shared library so its functions can be called from Python.
When import myextension runs for a compiled extension, what does Python look for first?
- A .py source file
- A Cargo.toml
- A shared library (.so, .pyd, .dll)
- A Dockerfile
Answer: A shared library (.so, .pyd, .dll). Python looks for a shared library (.so/.pyd/.dll), loads it, and calls its PyInit_ function.
Continue this course
- Previous: Automation & Scripting for DevOps and System Tasks
- Next: Architecture Patterns for Python Apps (MVC, Clean Architecture) — Structure large Python applications with proven architectural patterns
- Quick reference: Python cheat sheet