Add automatic constructor for classes

The languages with automatic constructors I showed still allow you to do stuff in the constructor, in addition to the automatic attribute assignment. And you can opt certain parameters out of the automatic assignment, (in TypeScript by omitting the accessibility modifier, and in Kotlin/Scala by omitting the var keyword").

You are right that some classes do not simply map their parameter arguments to attribute assignments, and those scenarios require the alternative notations. Personally, those are not very common, and I writing classes all the time as a primarily OOP GoF-esque developer.

I would wager if we analyzed a large sample of public codebases, we would find the average class scenario having a trivial constructor that would benefit from a shorthand constructor feature.

those languages have different idioms from Python, so it’s not self-evident that it makes sense to adopt this here.

Okay, but I think the burden of proof would be on you to argue why the benefits found in those languages would not be applicable when translated to Python idioms. I do OOP more or less the same way in all four of these languages. The idea of a class with a constructor doesn’t differ that much in languages, so we should expect the same considerations to transfer for the most part.

It would be much more compelling to actually do that analysis and show the value. Your intuition doesn’t match up with mine. That doesn’t mean you’re wrong, but it’s easier to convince people with evidence.

And I don’t know what “GoF-esque” means, sorry.

3 Likes

Gang of Four, one of the “Design Patterns” books that inspired a lot of the OOP trends. I just meant to say that I factor the majority of my codebases out of classes, so I write them a lot.

1 Like

I can believe that writing in a OOP-heavy style like that will enrich for this style of attribute-holder class. Python allows for an amalgamation of styles, though. I try to be a lot more functional in my code and I don’t tend to use classes so much, although I still use them depending on the task.

Python allows for an amalgamation of styles, though. I try to be a lot more functional in my code and I don’t tend to use classes so much,

Fair, but I don’t see any disadvantage to making the OOP style more ergonomic. It would be a net benefit, even if you personally would not make much of use of it. There are many who do make heavy use of classes in Python.

I don’t do much functional programming personally, but I would still appreciate Python choosing to enrich that experience for those who do by observing what primarily FP driven languages have come up with to make things more ergonomic and borrowing them.

Scala allows for both styles and really excels at both.

I actually wrote the __post_init__ of my implementation to handle the case where you want to do something but not to every attribute by having the init generator pick up the argument names defined in the post init function.

eg: if you write something roughly like:

@dataclass_like
class Ex:
    a: str
    b: Path
    
    def __post_init__(self, b: str | Path):
        self.b = Path(b)

You end up with:

class Ex:
    a: str
    b: Path

    def __init__(self, a: str, b: str | Path):
        self.a = a
        self.__post_init__(b=b)

    def __post_init__(self, b: str | Path):
        self.b = Path(b)

I do still generate __eq__ and __repr__ by default as I find those generally useful features to have on every python class and find those even more tedious to write than __init__ (It’s rare that I need hashability on these classes, if I need that I also need to make the class ‘immutable’).

1 Like

I was more thinking of cases where I don’t use a dataclass at all, because the style is more “take some parameters and construct an object” rather than storing the values themselves. I could use a dataclass with __post_init__ but I don’t really need the input parameters stored on the instance.

I was mostly mentioning the difference in style to gently push back on the idea that the “vast majority” of classes fit this pattern. I’m sure it’s fairly common, and maybe common enough to justify an addition here, but it’s always useful to consider how wide the space of python can be.

1 Like

to gently push back on the idea that the “vast majority” of classes fit this pattern.

Sure, I’ll concede that; I can only speak to the codebases I’ve worked on and seen anecdotally. Certainly the vast majority for me personally, and I would wager still a large portion for other users.

Yes that also works, I actually came up with a craftier version this morning where you didn’t need evolve_defaults as calling an instance of ClassBuilder without the positional only cls argument returned another ClassBuilder with whatever values you changed included in the new defaults. But then my power went out :frowning: .

I think when inheritance is involved? I would expect searching for a subclass __init__ to fall through to the parent if it is not defined and hence would expect the default to fall through to the implicit base class of object.

1 Like

Fair point @DavidCEllis, I can see why you might want to define ClassBuilder to turn off everything by default.

WRT processing of inputs (@jamestwebber, @dg-pb), the majority of classes I’ve seen can be neatly expressed using an autoinit in which you assign eg self.x = x and self._y = y, and post-processing in a __post_init__.
(And some classes just shouldn’t use an autoinit, because their init is doing something very specific.)
I understand your experience may be different.

For example I believe the example @DavidCEllis gave can currently be achieved as

@dataclass
class Ex:
    a: str
    b: Path | str
    
    def __post_init__(self):
        self.b = Path(self.b)

?

Or is the issue the construction was trying to solve something to do with Frozen=True or something like that?

There’s two elements to it, one is - as you say - for a frozen class this only sets the attribute once (iirc dataclasses’ implementation of frozen actually has to either assign directly to the __dict__ or use object.__setattr__ so you’d have to use that in post_init in that case).

The other idea is that the __init__ function accepts str | Path for b but the attribute is typed as Path only for the instance attribute. If you create an instance of your example and try to access instance.b.parent for example, mypy will complain that str doesn’t have a parent attribute. If you type b as Path it will complain that you’re using a string as input.[1]


  1. I’ll note that mypy can’t actually see the generated function - but if it could see it, it would be correct! ↩︎

1 Like

The example is illustrative, but in this case, shouldn’t it be typed as PathLike instead of str | Path :wink:

Or maybe str | PathLike, if typing is still broken in this regard …

Because that would break existing code. The existing defaults imply
specific behaviours and are conservative, eg around hashing. If
dataclass instances were hashable by default that would potentially
break everything that modified an instance after creation, or where the
attribute values themselves were unhashable.

That’s a no-fly thing - it simply cannot be done.

What could be done is an easy to access @frozen_dataclass or something
of that order (to be provided by the dataclasses module), where the
hash-relevant attributes were enumerated (maybe with some kind of
default) and type checked for being hashable themselves?

If you can sort out a concrete proposal for the above and get broad
agreement on the precise semantics I doubt there’s be much objection to
its existence or addition (via a PEP once you’ve got the specifics
broadly argued out here).

Nobody’s against some kind of hashable readonly data class thing. But
@dataclass itself does not quite mean that and that ship has sailed.
Build a better ship :slight_smile:

I agree, I wasn’t suggesting we change how dataclass works, but to add a new alternative to them, like a dedicated decorator or syntax notation.

1 Like

I think this was based on my example where I used str | Path. Technically the type in the __post_init__ method could have been str | Pathlike (I think as str doesn’t provide __fspath__ you still need the union) but the purpose was mainly to show how the type was copied to the __init__ and the narrowing from str | Path to Path going from __init__ parameter signature to instance attribute.