Descriptors & Properties: The Attribute Protocol
1 · The lesson
readEvery time you write obj.x, Python runs an algorithm. Most of the time it's "look in obj.__dict__, fall back to the class." But if the class attribute happens to be a descriptor, the algorithm hands control over to that descriptor's __get__ method — and the descriptor decides what obj.x returns.
That single hook is the foundation of @property, @classmethod, @staticmethod, method binding (the reason self works), functools.cached_property, every ORM field type (Django models.CharField, SQLAlchemy Column, Pydantic v1 fields), dataclasses.field(...), and a dozen other features you've used without realising. This lesson unpacks the protocol so those features stop being magic.
1. The Descriptor Protocol — One Sentence
A descriptor is any object whose class implements at least one of:
__get__(self, obj, objtype)— runs when the attribute is read.__set__(self, obj, value)— runs when the attribute is assigned.__delete__(self, obj)— runs when the attribute is deleted.
That's it. Define one of those, attach the instance as a class attribute on another class, and reads/writes of that attribute on instances run through your code.
class Const: def __get__(self, obj, objtype=None): return 42 class Foo: answer = Const() # descriptor instance, attached at class-body level f = Foo() print(f.answer) # 42 — Const.__get__ ran print(Foo.answer) # 42 — works on the class too
Three names to keep straight in __get__:
self— the descriptor instance (theConst()).obj— the instance the attribute is being accessed on (theFoo()), orNoneif accessed on the class.objtype— the class (Foo).
2. __set_name__ — Knowing Your Own Name
When you write x = Positive() inside a class body, the descriptor doesn't know it was named x. Python 3.6+ added __set_name__, which fires at class-creation time with the owner class and the attribute name:
class Tagged: def __set_name__(self, owner, name): print(f"attached to {owner.__name__} as {name!r}") self.name = name class Foo: a = Tagged() b = Tagged() # attached to Foo as 'a' # attached to Foo as 'b'
This is the modern foundation for any validating descriptor that needs to store per-instance state — you use name to namespace the value inside obj.__dict__ (see Section 4).
3. A Validating Descriptor — Positive
A descriptor that raises if you try to set a non-positive value:
class Positive: def __set_name__(self, owner, name): self._name = name # remember the attribute name def __get__(self, obj, objtype=None): if obj is None: # access on the class itself return self return obj.__dict__[self._name] def __set__(self, obj, value): if value <= 0: raise ValueError(f"{self._name} must be > 0, got {value}") obj.__dict__[self._name] = value class Order: quantity = Positive() price = Positive() def __init__(self, quantity, price): self.quantity = quantity # routes through Positive.__set__ self.price = price o = Order(3, 19.99) print(o.quantity, o.price) # 3 19.99 o.quantity = 5 # validated and stored # o.quantity = -1 # ValueError: quantity must be > 0, got -1 # Order(0, 10) # ValueError: quantity must be > 0, got 0
Two things to notice:
- State lives on the instance, not on the descriptor. We write into
obj.__dict__[self._name]. OnePositive()is shared by everyOrder; only the value is per-instance. obj is Nonecheck in__get__— when someone writesOrder.quantity(on the class, not an instance), return the descriptor itself so introspection still works.
4. Data vs Non-Data Descriptors — The Lookup Order
Python's attribute lookup has a specific priority:
1. Data descriptors on the class (have __set__ or __delete__) — win.
2. The instance __dict__.
3. Non-data descriptors on the class (only __get__) — lose to instance dict.
4. Class attributes (regular values).
class Loud: def __get__(self, obj, objtype=None): return "loud!" class Quiet: def __get__(self, obj, objtype=None): return "quiet" def __set__(self, obj, value): obj.__dict__["q"] = value class Foo: l = Loud() # non-data descriptor q = Quiet() # data descriptor f = Foo() f.__dict__["l"] = "shadowed" # instance dict wins over non-data descriptor f.__dict__["q"] = "shadowed" # data descriptor still wins print(f.l) # shadowed print(f.q) # quiet
This is why functions (non-data descriptors — they only define __get__) can be replaced per-instance by sticking something in instance.__dict__, but @property (a data descriptor) cannot.
5. Methods Are Descriptors
def-defined functions implement __get__. That's the entire mechanism behind method binding — self isn't a keyword, it's the descriptor protocol running:
class Foo: def hello(self): return f"hi from {self}" f = Foo() print(Foo.hello) # <function Foo.hello at 0x...> — the raw function print(f.hello) # <bound method Foo.hello of ...> — descriptor produced a bound method # What actually happens: import types manually_bound = Foo.hello.__get__(f, Foo) print(manually_bound) # <bound method Foo.hello of ...> print(manually_bound()) # hi from <__main__.Foo object>
Foo.hello.__get__(f, Foo) returns a types.MethodType that has f curried in as the first argument. Every method call goes through this — f.hello() is Foo.hello.__get__(f, Foo)().
Methods are non-data descriptors (no __set__), which is why you can monkey-patch a method on a specific instance by assigning to instance.method = ... — the instance dict wins.
6. @property IS a Descriptor
property is just a built-in descriptor class. The decorator form is convenience:
class C: @property def area(self): return self._r ** 2 * 3.14159
is roughly equivalent to:
class _AreaDescriptor: def __get__(self, obj, objtype=None): if obj is None: return self return obj._r ** 2 * 3.14159 class C: area = _AreaDescriptor()
The real property is a single class that handles fget, fset, and fdel via __init__ arguments — @property.setter builds a new property object with the setter slot filled in. Understanding this is what makes "why does my property setter not work after I redefine the getter in a subclass?" intuitive instead of mysterious.
@classmethod and @staticmethod are also descriptors. classmethod.__get__ returns a bound method with the class curried in instead of the instance; staticmethod.__get__ returns the underlying function unbound. You almost never need to know this — but it explains why all three "method types" can sit at the class-body level and behave consistently at access time.
7. A Real-World Descriptor — Lazy
Compute a value once on first access, then cache it on the instance for every subsequent read. Pattern: write into obj.__dict__ from __get__. Because regular attribute lookup checks the instance dict before non-data descriptors, the cached value is found there directly on the next call — your __get__ never runs again.
class Lazy: def __init__(self, func): self.func = func def __set_name__(self, owner, name): self._name = name def __get__(self, obj, objtype=None): if obj is None: return self value = self.func(obj) obj.__dict__[self._name] = value # cache — shadows the descriptor next time return value class Dataset: def __init__(self, rows): self.rows = rows @Lazy def summary(self): print("computing summary (expensive)...") return { "count": len(self.rows), "total": sum(self.rows), } d = Dataset([1, 2, 3, 4, 5]) print(d.summary) # computing summary (expensive)... {'count': 5, 'total': 15} print(d.summary) # {'count': 5, 'total': 15} — no recompute print(d.summary) # ditto
The stdlib already ships this — use functools.cached_property:
from functools import cached_property class Dataset: def __init__(self, rows): self.rows = rows @cached_property def summary(self): return {"count": len(self.rows), "total": sum(self.rows)}
Always prefer the stdlib one. It handles edge cases (slotted classes, thread safety hooks, __set_name__ plumbing) you don't want to reinvent. Writing your own Lazy is a teaching exercise; importing cached_property is the answer in production.
8. When to Write Your Own Descriptor
The honest rule:
- One attribute, one class, simple logic →
@property. - The same validation logic on many attributes across many classes → descriptor.
Five @propertys with copy-pasted validation are the smell. One Typed(int) descriptor used on quantity, count, index, age across Order, Inventory, Page, User is the win — you write the validation once and apply it everywhere.
Real-world descriptor users:
- Django ORM —
models.CharField(max_length=100)returns a descriptor that mediates DB column access. - SQLAlchemy —
Column(...)andrelationship(...)are descriptors that proxy to the session. - Pydantic v1 —
ModelFielddescriptors run validators on set. dataclasses.field(...)— produces a descriptor (since 3.10) that handles defaults, factory functions, and thedefault_factoryprotocol.
You'll mostly configure these. You'll occasionally write a small one for a project-specific validation rule. Anything beyond that is usually a sign you're building a framework.
9. Descriptors and dataclasses
Since Python 3.10, dataclass fields cooperate with the descriptor protocol — if you put a descriptor in a field annotation, the dataclass-generated __init__ will go through it:
from dataclasses import dataclass class NonEmpty: def __set_name__(self, owner, name): self._name = name def __get__(self, obj, objtype=None): if obj is None: return self return obj.__dict__[self._name] def __set__(self, obj, value): if not value: raise ValueError(f"{self._name} cannot be empty") obj.__dict__[self._name] = value @dataclass class User: name: str = NonEmpty() email: str = NonEmpty() u = User("Linus", "linus@example.com") print(u) # User(name='Linus', email='linus@example.com') # User("", "x@y") # ValueError: name cannot be empty
A descriptor lets you compose validation cleanly with dataclasses — much nicer than __post_init__ for the simple "this field must satisfy a rule" case.
Common Mistakes
1. Storing per-instance state on the descriptor itself — the classic bug
class Bad: def __get__(self, obj, objtype=None): return self.value def __set__(self, obj, value): self.value = value # WRONG — written onto the descriptor instance, # shared across every owning instance! class Foo: x = Bad() a, b = Foo(), Foo() a.x = 1 b.x = 2 print(a.x) # 2 — they share!
The descriptor is one object held as a class attribute. Store the value in obj.__dict__[self._name] instead — that's per-instance.
2. Confusing obj and objtype in __get__
obj is the instance (or None if accessed on the class). objtype is the class. Mixing them up gives you methods that return class-level data when you wanted instance-level (or vice versa). Always handle the obj is None case explicitly — return self so introspection on the class still works.
3. Forgetting __set_name__ and manually passing names
class Field: def __init__(self, name): # fragile — must match the attribute name self.name = name class Foo: x = Field("x") # repeated yourself y = Field("y") # easy to typo
__set_name__ gets the name for free at class-creation time. Use it. Pre-3.6 code passes names manually because it had no choice; new code shouldn't.
4. Writing a descriptor when a @property would do
A descriptor used in exactly one place, by one attribute, on one class, is over-engineering. @property is shorter, more idiomatic, and easier for the next person to read. Promote to a descriptor only when you'd be writing the same logic for the third time.
5. Expecting descriptors to fire on dict-style access
class Foo: x = SomeDescriptor() f = Foo() f.x # descriptor fires f.__dict__["x"] = 42 # bypasses the descriptor — direct dict write print(f.__dict__) # {'x': 42}
setup added so this can run · defines SomeDescriptor
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def SomeDescriptor(*_a, **_kw): print('-> SomeDescriptor() called') return _AutoMock('SomeDescriptor()')
Descriptors hook into normal attribute access (f.x = ...). Reaching directly into __dict__ skips them. This is usually a feature (you use it from inside the descriptor itself to store state), but it's a footgun when test code or setattr magic does it accidentally.
🎯 Your Turn — Build a Typed Descriptor
Build Typed(expected_type) — a descriptor that enforces a type on every set, raising TypeError with a clear message that includes the attribute name. Requirements:
name = Typed(str)andage = Typed(int)should work as class attributes.- Setting a value of the wrong type raises
TypeError: <attr_name> must be <type>, got <actual_type>. - Setting a value of the correct type stores it per-instance (no shared state between instances).
- Reading returns the stored value. Reading before any set should raise
AttributeError(Python's default for missing attributes). - Use
__set_name__so you don't have to pass the name in manually.
class Typed: def __init__(self, expected_type): # TODO 1: store expected_type ... def __set_name__(self, owner, name): # TODO 2: remember the attribute name ... def __get__(self, obj, objtype=None): # TODO 3: handle obj is None (return self) # TODO 4: pull the value from obj.__dict__ ... def __set__(self, obj, value): # TODO 5: isinstance check; raise TypeError if wrong # TODO 6: store in obj.__dict__[self._name] ... class Person: name = Typed(str) age = Typed(int) def __init__(self, name, age): self.name = name self.age = age p = Person("Linus", 28) print(p.name, p.age) # Person(123, 28) # TypeError: name must be <class 'str'>, got <class 'int'>
Hint 1 — Per-instance storage via obj.__dict__
Inside __set__, write obj.__dict__[self._name] = value. Inside __get__, read it back the same way. Never write self.value = ... on the descriptor — that's the shared-state bug from Common Mistake #1.
Hint 2 — Using __set_name__ for the attribute name
Python calls __set_name__(self, owner, name) automatically when the class body finishes. Save name as self._name so your error messages and the __dict__ key can use it. No need for the user to pass the name as an argument.
Show full solution
class Typed: def __init__(self, expected_type): self.expected_type = expected_type def __set_name__(self, owner, name): self._name = name def __get__(self, obj, objtype=None): if obj is None: return self # access on the class try: return obj.__dict__[self._name] except KeyError: raise AttributeError( f"{type(obj).__name__!r} object has no attribute {self._name!r}" ) from None def __set__(self, obj, value): if not isinstance(value, self.expected_type): raise TypeError( f"{self._name} must be {self.expected_type}, got {type(value)}" ) obj.__dict__[self._name] = value class Person: name = Typed(str) age = Typed(int) def __init__(self, name, age): self.name = name self.age = age # Happy path p = Person("Linus", 28) print(p.name, p.age) # Linus 28 p.age = 29 print(p.age) # 29 # Type errors try: Person(123, 28) except TypeError as e: print(e) # name must be <class 'str'>, got <class 'int'> try: p.age = "old" except TypeError as e: print(e) # age must be <class 'int'>, got <class 'str'> # Per-instance isolation — no shared state p2 = Person("Ada", 36) print(p.name, p2.name) # Linus Ada — independent
A few production-grade refinements you'd add in real code:
boolsubclassesintin Python, soTyped(int)acceptsTrue/False. If you want to reject bools, addand not isinstance(value, bool).- Subclass acceptance —
isinstanceaccepts subclasses by design.Typed(Animal)will accept anyDog. Usually what you want; occasionally not, in which case usetype(value) is self.expected_type. - Tuples of types —
isinstanceaccepts(int, float). Make the constructor accept a tuple for "any of these":Typed((int, float))works without changing a line. - Default values — supply a
default=argument in__init__and return it from__get__when the key is missing, instead of raisingAttributeError.
You've just built the core of every typed-field library out there — Pydantic's Field, attrs' attr.ib(type=...), Django's models.CharField. Same protocol, more bells and whistles.
What You Learned
- A descriptor is any object whose class defines
__get__,__set__, or__delete__. It mediates attribute access on instances of whatever class it's a class-attribute of. __set_name__(self, owner, name)(3.6+) gives the descriptor its own attribute name at class-creation time — the modern foundation for any validating descriptor.- Data descriptors (have
__set__) win over instance__dict__. Non-data descriptors (only__get__) lose. That's why methods can be monkey-patched on instances but@propertycannot. - Functions are non-data descriptors — their
__get__returns a bound method withselfcurried in. That's the entire mechanism behind method calls. @property,@classmethod,@staticmethodare all built-in descriptor classes. The decorator form is sugar.functools.cached_propertyis the stdlib descriptor for "compute once, cache on the instance". Always prefer it over hand-rolledLazy.- Store per-instance state in
obj.__dict__[self._name], never onself(the descriptor itself is shared across every owning instance). - Write your own descriptor only when the same logic is repeated across many attributes across many classes. Otherwise
@propertyis enough.
Next: Context Managers — the with protocol, __enter__/__exit__, and contextlib for resource-safe code that cleans up on every path.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.