PythonMastery
advanced 22 min read · lesson 2 of 9 in Python Advanced

Descriptors & Properties: The Attribute Protocol

1 · The lesson

read

Every time you write obj.x, Python runs an algorithm. Most of the time it's "look in obj.__dict__, fall back to the class." But if the class attribute happens to be a descriptor, the algorithm hands control over to that descriptor's __get__ method — and the descriptor decides what obj.x returns.

That single hook is the foundation of @property, @classmethod, @staticmethod, method binding (the reason self works), functools.cached_property, every ORM field type (Django models.CharField, SQLAlchemy Column, Pydantic v1 fields), dataclasses.field(...), and a dozen other features you've used without realising. This lesson unpacks the protocol so those features stop being magic.


1. The Descriptor Protocol — One Sentence

A descriptor is any object whose class implements at least one of:

  • __get__(self, obj, objtype) — runs when the attribute is read.
  • __set__(self, obj, value) — runs when the attribute is assigned.
  • __delete__(self, obj) — runs when the attribute is deleted.

That's it. Define one of those, attach the instance as a class attribute on another class, and reads/writes of that attribute on instances run through your code.

python
class Const:
    def __get__(self, obj, objtype=None):
        return 42

class Foo:
    answer = Const()            # descriptor instance, attached at class-body level

f = Foo()
print(f.answer)                 # 42           — Const.__get__ ran
print(Foo.answer)               # 42           — works on the class too

Three names to keep straight in __get__:

  • self — the descriptor instance (the Const()).
  • obj — the instance the attribute is being accessed on (the Foo()), or None if accessed on the class.
  • objtype — the class (Foo).

2. __set_name__ — Knowing Your Own Name

When you write x = Positive() inside a class body, the descriptor doesn't know it was named x. Python 3.6+ added __set_name__, which fires at class-creation time with the owner class and the attribute name:

python
class Tagged:
    def __set_name__(self, owner, name):
        print(f"attached to {owner.__name__} as {name!r}")
        self.name = name

class Foo:
    a = Tagged()
    b = Tagged()
# attached to Foo as 'a'
# attached to Foo as 'b'

This is the modern foundation for any validating descriptor that needs to store per-instance state — you use name to namespace the value inside obj.__dict__ (see Section 4).


3. A Validating Descriptor — Positive

A descriptor that raises if you try to set a non-positive value:

python
class Positive:
    def __set_name__(self, owner, name):
        self._name = name                       # remember the attribute name

    def __get__(self, obj, objtype=None):
        if obj is None:                         # access on the class itself
            return self
        return obj.__dict__[self._name]

    def __set__(self, obj, value):
        if value <= 0:
            raise ValueError(f"{self._name} must be > 0, got {value}")
        obj.__dict__[self._name] = value


class Order:
    quantity = Positive()
    price = Positive()

    def __init__(self, quantity, price):
        self.quantity = quantity                # routes through Positive.__set__
        self.price = price


o = Order(3, 19.99)
print(o.quantity, o.price)      # 3 19.99
o.quantity = 5                  # validated and stored
# o.quantity = -1               # ValueError: quantity must be > 0, got -1
# Order(0, 10)                  # ValueError: quantity must be > 0, got 0

Two things to notice:

  • State lives on the instance, not on the descriptor. We write into obj.__dict__[self._name]. One Positive() is shared by every Order; only the value is per-instance.
  • obj is None check in __get__ — when someone writes Order.quantity (on the class, not an instance), return the descriptor itself so introspection still works.

4. Data vs Non-Data Descriptors — The Lookup Order

Python's attribute lookup has a specific priority:

1. Data descriptors on the class (have __set__ or __delete__) — win.
2. The instance __dict__.
3. Non-data descriptors on the class (only __get__) — lose to instance dict.
4. Class attributes (regular values).

python
class Loud:
    def __get__(self, obj, objtype=None):
        return "loud!"

class Quiet:
    def __get__(self, obj, objtype=None):
        return "quiet"
    def __set__(self, obj, value):
        obj.__dict__["q"] = value

class Foo:
    l = Loud()                  # non-data descriptor
    q = Quiet()                 # data descriptor

f = Foo()
f.__dict__["l"] = "shadowed"    # instance dict wins over non-data descriptor
f.__dict__["q"] = "shadowed"    # data descriptor still wins

print(f.l)                      # shadowed
print(f.q)                      # quiet

This is why functions (non-data descriptors — they only define __get__) can be replaced per-instance by sticking something in instance.__dict__, but @property (a data descriptor) cannot.


5. Methods Are Descriptors

def-defined functions implement __get__. That's the entire mechanism behind method binding — self isn't a keyword, it's the descriptor protocol running:

python
class Foo:
    def hello(self):
        return f"hi from {self}"

f = Foo()
print(Foo.hello)                # <function Foo.hello at 0x...>     — the raw function
print(f.hello)                  # <bound method Foo.hello of ...>   — descriptor produced a bound method

# What actually happens:
import types
manually_bound = Foo.hello.__get__(f, Foo)
print(manually_bound)           # <bound method Foo.hello of ...>
print(manually_bound())         # hi from <__main__.Foo object>

Foo.hello.__get__(f, Foo) returns a types.MethodType that has f curried in as the first argument. Every method call goes through this — f.hello() is Foo.hello.__get__(f, Foo)().

Methods are non-data descriptors (no __set__), which is why you can monkey-patch a method on a specific instance by assigning to instance.method = ... — the instance dict wins.


6. @property IS a Descriptor

property is just a built-in descriptor class. The decorator form is convenience:

python
class C:
    @property
    def area(self):
        return self._r ** 2 * 3.14159

is roughly equivalent to:

python
class _AreaDescriptor:
    def __get__(self, obj, objtype=None):
        if obj is None:
            return self
        return obj._r ** 2 * 3.14159

class C:
    area = _AreaDescriptor()

The real property is a single class that handles fget, fset, and fdel via __init__ arguments — @property.setter builds a new property object with the setter slot filled in. Understanding this is what makes "why does my property setter not work after I redefine the getter in a subclass?" intuitive instead of mysterious.

@classmethod and @staticmethod are also descriptors. classmethod.__get__ returns a bound method with the class curried in instead of the instance; staticmethod.__get__ returns the underlying function unbound. You almost never need to know this — but it explains why all three "method types" can sit at the class-body level and behave consistently at access time.


7. A Real-World Descriptor — Lazy

Compute a value once on first access, then cache it on the instance for every subsequent read. Pattern: write into obj.__dict__ from __get__. Because regular attribute lookup checks the instance dict before non-data descriptors, the cached value is found there directly on the next call — your __get__ never runs again.

python
class Lazy:
    def __init__(self, func):
        self.func = func

    def __set_name__(self, owner, name):
        self._name = name

    def __get__(self, obj, objtype=None):
        if obj is None:
            return self
        value = self.func(obj)
        obj.__dict__[self._name] = value        # cache — shadows the descriptor next time
        return value


class Dataset:
    def __init__(self, rows):
        self.rows = rows

    @Lazy
    def summary(self):
        print("computing summary (expensive)...")
        return {
            "count": len(self.rows),
            "total": sum(self.rows),
        }


d = Dataset([1, 2, 3, 4, 5])
print(d.summary)                # computing summary (expensive)... {'count': 5, 'total': 15}
print(d.summary)                # {'count': 5, 'total': 15}        — no recompute
print(d.summary)                # ditto

The stdlib already ships this — use functools.cached_property:

python
from functools import cached_property

class Dataset:
    def __init__(self, rows):
        self.rows = rows

    @cached_property
    def summary(self):
        return {"count": len(self.rows), "total": sum(self.rows)}

Always prefer the stdlib one. It handles edge cases (slotted classes, thread safety hooks, __set_name__ plumbing) you don't want to reinvent. Writing your own Lazy is a teaching exercise; importing cached_property is the answer in production.


8. When to Write Your Own Descriptor

The honest rule:

  • One attribute, one class, simple logic → @property.
  • The same validation logic on many attributes across many classes → descriptor.

Five @propertys with copy-pasted validation are the smell. One Typed(int) descriptor used on quantity, count, index, age across Order, Inventory, Page, User is the win — you write the validation once and apply it everywhere.

Real-world descriptor users:

  • Django ORM — models.CharField(max_length=100) returns a descriptor that mediates DB column access.
  • SQLAlchemy — Column(...) and relationship(...) are descriptors that proxy to the session.
  • Pydantic v1 — ModelField descriptors run validators on set.
  • dataclasses.field(...) — produces a descriptor (since 3.10) that handles defaults, factory functions, and the default_factory protocol.

You'll mostly configure these. You'll occasionally write a small one for a project-specific validation rule. Anything beyond that is usually a sign you're building a framework.


9. Descriptors and dataclasses

Since Python 3.10, dataclass fields cooperate with the descriptor protocol — if you put a descriptor in a field annotation, the dataclass-generated __init__ will go through it:

python
from dataclasses import dataclass

class NonEmpty:
    def __set_name__(self, owner, name):
        self._name = name
    def __get__(self, obj, objtype=None):
        if obj is None: return self
        return obj.__dict__[self._name]
    def __set__(self, obj, value):
        if not value:
            raise ValueError(f"{self._name} cannot be empty")
        obj.__dict__[self._name] = value


@dataclass
class User:
    name: str = NonEmpty()
    email: str = NonEmpty()


u = User("Linus", "linus@example.com")
print(u)                        # User(name='Linus', email='linus@example.com')
# User("", "x@y")               # ValueError: name cannot be empty

A descriptor lets you compose validation cleanly with dataclasses — much nicer than __post_init__ for the simple "this field must satisfy a rule" case.


Common Mistakes

1. Storing per-instance state on the descriptor itself — the classic bug

python
class Bad:
    def __get__(self, obj, objtype=None):
        return self.value
    def __set__(self, obj, value):
        self.value = value          # WRONG — written onto the descriptor instance,
                                    #         shared across every owning instance!

class Foo:
    x = Bad()

a, b = Foo(), Foo()
a.x = 1
b.x = 2
print(a.x)                          # 2     — they share!

The descriptor is one object held as a class attribute. Store the value in obj.__dict__[self._name] instead — that's per-instance.

2. Confusing obj and objtype in __get__

obj is the instance (or None if accessed on the class). objtype is the class. Mixing them up gives you methods that return class-level data when you wanted instance-level (or vice versa). Always handle the obj is None case explicitly — return self so introspection on the class still works.

3. Forgetting __set_name__ and manually passing names

python
class Field:
    def __init__(self, name):       # fragile — must match the attribute name
        self.name = name

class Foo:
    x = Field("x")                  # repeated yourself
    y = Field("y")                  # easy to typo

__set_name__ gets the name for free at class-creation time. Use it. Pre-3.6 code passes names manually because it had no choice; new code shouldn't.

4. Writing a descriptor when a @property would do

A descriptor used in exactly one place, by one attribute, on one class, is over-engineering. @property is shorter, more idiomatic, and easier for the next person to read. Promote to a descriptor only when you'd be writing the same logic for the third time.

5. Expecting descriptors to fire on dict-style access

python
class Foo:
    x = SomeDescriptor()

f = Foo()
f.x                                 # descriptor fires
f.__dict__["x"] = 42                # bypasses the descriptor — direct dict write
print(f.__dict__)                   # {'x': 42}
+ setup added so this can run · defines SomeDescriptor
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def SomeDescriptor(*_a, **_kw):
    print('-> SomeDescriptor() called')
    return _AutoMock('SomeDescriptor()')

Descriptors hook into normal attribute access (f.x = ...). Reaching directly into __dict__ skips them. This is usually a feature (you use it from inside the descriptor itself to store state), but it's a footgun when test code or setattr magic does it accidentally.


🎯 Your Turn — Build a Typed Descriptor

Build Typed(expected_type) — a descriptor that enforces a type on every set, raising TypeError with a clear message that includes the attribute name. Requirements:

  • name = Typed(str) and age = Typed(int) should work as class attributes.
  • Setting a value of the wrong type raises TypeError: <attr_name> must be <type>, got <actual_type>.
  • Setting a value of the correct type stores it per-instance (no shared state between instances).
  • Reading returns the stored value. Reading before any set should raise AttributeError (Python's default for missing attributes).
  • Use __set_name__ so you don't have to pass the name in manually.
python
class Typed:
    def __init__(self, expected_type):
        # TODO 1: store expected_type
        ...

    def __set_name__(self, owner, name):
        # TODO 2: remember the attribute name
        ...

    def __get__(self, obj, objtype=None):
        # TODO 3: handle obj is None (return self)
        # TODO 4: pull the value from obj.__dict__
        ...

    def __set__(self, obj, value):
        # TODO 5: isinstance check; raise TypeError if wrong
        # TODO 6: store in obj.__dict__[self._name]
        ...


class Person:
    name = Typed(str)
    age = Typed(int)

    def __init__(self, name, age):
        self.name = name
        self.age = age


p = Person("Linus", 28)
print(p.name, p.age)
# Person(123, 28)               # TypeError: name must be <class 'str'>, got <class 'int'>
Hint 1 — Per-instance storage via obj.__dict__ Inside __set__, write obj.__dict__[self._name] = value. Inside __get__, read it back the same way. Never write self.value = ... on the descriptor — that's the shared-state bug from Common Mistake #1.
Hint 2 — Using __set_name__ for the attribute name Python calls __set_name__(self, owner, name) automatically when the class body finishes. Save name as self._name so your error messages and the __dict__ key can use it. No need for the user to pass the name as an argument.
Show full solution
python
class Typed:
    def __init__(self, expected_type):
        self.expected_type = expected_type

    def __set_name__(self, owner, name):
        self._name = name

    def __get__(self, obj, objtype=None):
        if obj is None:
            return self                                 # access on the class
        try:
            return obj.__dict__[self._name]
        except KeyError:
            raise AttributeError(
                f"{type(obj).__name__!r} object has no attribute {self._name!r}"
            ) from None

    def __set__(self, obj, value):
        if not isinstance(value, self.expected_type):
            raise TypeError(
                f"{self._name} must be {self.expected_type}, got {type(value)}"
            )
        obj.__dict__[self._name] = value


class Person:
    name = Typed(str)
    age = Typed(int)

    def __init__(self, name, age):
        self.name = name
        self.age = age


# Happy path
p = Person("Linus", 28)
print(p.name, p.age)            # Linus 28
p.age = 29
print(p.age)                    # 29

# Type errors
try:
    Person(123, 28)
except TypeError as e:
    print(e)                    # name must be <class 'str'>, got <class 'int'>

try:
    p.age = "old"
except TypeError as e:
    print(e)                    # age must be <class 'int'>, got <class 'str'>

# Per-instance isolation — no shared state
p2 = Person("Ada", 36)
print(p.name, p2.name)          # Linus Ada            — independent

A few production-grade refinements you'd add in real code:

  • bool subclasses int in Python, so Typed(int) accepts True/False. If you want to reject bools, add and not isinstance(value, bool).
  • Subclass acceptance — isinstance accepts subclasses by design. Typed(Animal) will accept any Dog. Usually what you want; occasionally not, in which case use type(value) is self.expected_type.
  • Tuples of types — isinstance accepts (int, float). Make the constructor accept a tuple for "any of these": Typed((int, float)) works without changing a line.
  • Default values — supply a default= argument in __init__ and return it from __get__ when the key is missing, instead of raising AttributeError.

You've just built the core of every typed-field library out there — Pydantic's Field, attrs' attr.ib(type=...), Django's models.CharField. Same protocol, more bells and whistles.


What You Learned

  • A descriptor is any object whose class defines __get__, __set__, or __delete__. It mediates attribute access on instances of whatever class it's a class-attribute of.
  • __set_name__(self, owner, name) (3.6+) gives the descriptor its own attribute name at class-creation time — the modern foundation for any validating descriptor.
  • Data descriptors (have __set__) win over instance __dict__. Non-data descriptors (only __get__) lose. That's why methods can be monkey-patched on instances but @property cannot.
  • Functions are non-data descriptors — their __get__ returns a bound method with self curried in. That's the entire mechanism behind method calls.
  • @property, @classmethod, @staticmethod are all built-in descriptor classes. The decorator form is sugar.
  • functools.cached_property is the stdlib descriptor for "compute once, cache on the instance". Always prefer it over hand-rolled Lazy.
  • Store per-instance state in obj.__dict__[self._name], never on self (the descriptor itself is shared across every owning instance).
  • Write your own descriptor only when the same logic is repeated across many attributes across many classes. Otherwise @property is enough.

Next: Context Managers — the with protocol, __enter__/__exit__, and contextlib for resource-safe code that cleans up on every path.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.