Sunday, 4 October 2026

Python Bound Methods Optimization

We know Python is slow when compared to environments with massively optimized JIT's like the JVM or node.js (using more than one JIT, PGO, etc), but for an interpreted language with no JIT one could say that Python feels fast, particularly if we are aware of how, due to its enormous dynamism, basic actions involve rather complex steps. We saw here and here that calling a function (that at a semantic level means searching for a __call__ method in its type) is done in a more straightforward way.

We're going to look into some optimizations in this post and the next one. A key element is that these optimizations do not compromise language semantics. Under certain common circumstances the optimization will shortcut the way an action works based on the language semantics (like invoking a descriptor to create a bound method or searching the MRO to retrieve an attribute), but if we perform any changes that will render different the output of the optimization from the expected output based on language semantics, the optimization is reverted. Better explained by a GPT:

CPython may combine descriptor handling, instance storage, MRO lookup, a type-level lookup cache, per bytecode inline caches, version guards, and call paths that avoid temporary bound-method objects.
The central principle is:
The runtime must preserve every observable consequence of Python’s semantics, but it need not perform unobservable intermediate work.

Method invokation, bound-method elision
Invoking a method means looking up an attribute. That look up will find a function in a class, and as functions are descriptors the descriptor protocol will run returning a bound method (with __self__ and __func__ attributes). Then the bound method is called (semantically by looking up its __call__ attribute, but we know, from previous linked posts, that that gets optimized). The thing is that in most cases that bound method creation can be skipped. Let me explain:


class Person:
    def __init__(self, name):
        self.name = name

    def greet(self):
        return f"Hello, my name is {self.name}."    

p1 = Person("Iyan")
b_m = p1.greet
print(type(b_m ))  
print(b_m.__func__)
print(b_m.__self__)
# 
# 
# <__main__.Person object at 0x74ef724286e0>

print(b_m("Francois"))
# Hello Francois my name is Iyan.


In this case, the result of attribute access becomes visible to Python code (as we store it in the b_m variable for later use). For an ordinary Python function found in the class, the descriptor protocol produces a bound method containing __self__ and __func__, and calling that bound method is conceptually equivalent to: b_m.__func__(b_m.__self__, "Francois")

But most times we just invoke a method directly, like this:


p1.greet("Francois")

Creating a temporary method object merely to unpack those two references would be wasteful. CPython could instead retain the function and receiver separately and invoke: Person.say_hi(p1, "Xuan"). And basically that's what modern CPython does.

If we look at the bytecodes generated for this second case we'll see:


 23           LOAD_FAST_BORROW         0 (p)
              LOAD_ATTR                3 (greet + NULL|self)
              LOAD_CONST               1 ('Francois')
              CALL                     1
              POP_TOP
              LOAD_CONST               2 (None)
              RETURN_VALUE

That NULL|self tells the lookup algorithm to look up the function and put it in the stack (CPython is a stack based VM) along with the self value (the receiver, the object on which we've done the look up). After that, other parameters to be passed to the function are pushed in the stack and the function is called.

In the next post we'll see the optimizations used for attribute lookup