PEP 802: Display Syntax for the Empty Set

Nice comparison, but both increase the learning curve. For beginners, some constructs may be difficult to interpret without prior exposure. Even today, I had to look up what they actually mean (which is which).

In general, we should avoid unnecessarily increasing the learning curve. Both examples rely on prior knowledge of the language’s conventions and syntax, which becomes automatic with experience.


This shows 4.3 million results today.

1 Like

This is all just bikeshedding. At the end of the day, the PEP author needs to just make a decision - there’s never going to be consensus.

The real problem with the PEP is still that the motivation is weak, and it (IMO) fails to make the case that we need to do anything. The lack of consensus over spelling also weakens the case, but there’s no sign that will change, so arguing over it simply emphasises that people can’t agree on what looks right.

26 Likes

I only resurrected this 15-day-dormant thread because I noticed something no one else had :sweat_smile: Dozens of people worked on creating clunky UTF-16 but two geniuses came up with beautiful UTF-8; which one won? The point being just because only one person noticed the congruence with positional-only arguments doesn’t discredit it. /

I agree that there may not be consensus on the syntax and that courage to commit to one should be taken. I do like the motivation though. (), [], {}, {/} besties.

Also I have opened a PR for a slight change in the wording of the mention of my post in the PEP. I perhaps should’ve made it clearer that I’m in favor of allowing spaces like { / }.

I’ve tried the reference implementation and it’s looking good.

1 Like

No, but it doesn’t add any weight to it either. As Paul said, this proposal has two critical issues:

  1. Why even do this?
  2. Nobody can agree on a spelling.

Arguing about one specific spelling is more proof of the second problem, and doesn’t solve the (more important) first problem. So by all means, come up with reasons why YOUR particular preferred spelling is best; just understand that, the more you argue about spelling without giving motivation, the weaker you make this proposal overall.

4 Likes

For me it’s about giving Python an ascii ideograph for the empty set, suitable for both writing it in programs, and reading it in repr and str strings. It makes set in league with list dict and tuple. It looks like Ø, but using curly braces instead of other ([< etc because that’s what Python’s sets use.

It topped the poll (not that polls are instructive though, just inquisitive) and there’ll always be other people with different taste. I find rejected alternatives compelling. You and 2% of the poll can continue to use {*()} if you want, but {/} feels right being more declarative than {*()} and set() which seem imperative.


>>> empty_tuple =  ()
>>> empty_list  =  []
>>> empty_dict  =  {}
>>> empty_set   = {/}
>>> 
>>> empties = [empty_tuple, empty_list, empty_dict, empty_set]
>>> 
>>> for empty in empties:
...     print(f"{str(empty) = !s:>3} {repr(empty) = !s:>3}")
...     
str(empty) =  () repr(empty) =  ()
str(empty) =  [] repr(empty) =  []
str(empty) =  {} repr(empty) =  {}
str(empty) = {/} repr(empty) = {/}

Yummers

Yes, I know. I HAVE read the thread (at least, most of it; it’s getting pretty long at this point). The point is not “there is no reason whatsoever to do this”, the point is “that justification is not strong”. Restating the same weak reason for doing it is simply not sufficient.

Polls tell you a whole lot of nothing. Go through the history of successful proposals for Python and see how many of them cite polls; off the top of my head, I can only think of one, and the winning entry wasn’t what ended up happening.

We get it. You like this specific syntax. But continuing to debate syntax choices isn’t going to solve the underlying and very fundamental problems. Figure out a much stronger justification or give up on the thread. I’m not going to tell you what my preferred syntax is, and I’m not going to tell you what I actually do in my actual code, because it’s not relevant.

4 Likes

Reading this whole thread gives me the impression that the community isn’t going to reach a consensus on whether having a new literal for the empty set outweighs the costs, nor on which syntax is the “best” syntax if it does. People are occasionally throwing in their two cents, but that only seems to further establish that everyone has different strong opinions on this topic without universally compelling arguments on why their opinion should be considered superior to others. (That is not to say any new arguments for or against a bare empty set literal should be shut down. I simply feel that subjective opinions will contribute little towards reaching a consensus at this point.)

Because most of the people in favor of the status quo seem to be of the opinion that the addition of a somewhat unconventional and inconsistent syntax for an empty set literal would result in valid but only marginal gains, the discussion seems to be shifting towards a more general syntax that naturally implies one of the proposed empty set literals.

I’m seeing two distinct proposals in the thread at the moment.

  1. / to indicate absence (link 1) (link 2)
    • / inserted in any context that normally takes zero or more values is treated equivalent to not giving a value at that position at all. (see original text for examples; I’m not quoting them here due to length and because they haven’t been completely fleshed out or thoroughly checked for inconsistencies)
      • Makes constructing variable-length containers without separate insertions/deletions or iteration tools easier, and potentially allows for new match-case syntax to match absence of dictionary key
      • One immediate concern I’m seeing is that new learners could easily confuse this syntax with the existing positional-only indicator /.
    • Naturally implies {/} as empty set literal, although the choice of the character / is somewhat arbitrary compared to the following proposal
  2. Null Literal Unpacking
    • Use bare * and ** inside container characters ([], (), {}) to indicate emptiness
    • Has consistency with existing unpacking syntax: [*a, *b], {*a, *b}, {**x, **y}
    • Naturally implies {*} as empty set literal and {**} as empty dict literal

While I feel both of these proposals have merit enough to further develop, I’d suggest that they be moved to their respective discussions so that they may be fleshed out without being strongly tethered to the original topic of this thread.

4 Likes

+1
Having read this exchange and the PY3K mailing list thread on sets, it seems that the actual decision that caused this topic was the addition of the literal set syntax ({…, …}) in Python 3, and from what I can sense here, it’s likely that the same people opposing any literal syntax for an empty set now would have opposed that decision as well.

So the underlying question then, is purely one of consistency in the language itself: Are sets meant to be seen as “first-class” data structures in the Python programming language, in the same way that tuples, lists and dicts are?

To explain the idea behind “first-class”, I’m leaning into how the average beginner sees Python code: A programming language that is easy to understand because it reads like English, which also happens to be the dominant language used in published computer science concepts worldwide (and which eventually includes writing sets literally with curly braces). From that perspective, I definitely think it’s a win to have added the syntax, given how it leverages an intuition people already know from set theory, as a means to suggest that the use of efficient data structures for membership testing in dynamic code is so well recommended by the language itself, that is in fact part of the language’s grammar as a literal set (and thus implied to be ‘first-class’, just like tuples, lists and dicts with their literal syntax).

So, given that the literal set syntax made its way in, why not lean into the consistency expected by that same intuition of importing set theory ideas, and add the {/} syntax for completeness sake?

IMHO, to not do so is to suggest that the initial addition of the literal syntax was a design mistake, which would at least explain the unexpected result of repr({1} & {2}) == 'set()' to a beginner. It’s a valid path that would only need a docs change to discourage literal set use, but it’s also not a desirable one for those who actually use sets frequently in their code.

Zooming out, I think there were only ever 3 possible timelines:

  1. {:} and {/} were added together
  2. {:} was added first for dicts, so {} remains for sets
  3. {} was added first for dicts, so {/} remains for sets

We’re already on timeline 3, so the obvious path is to approve the addition of {/}.

Generalizing / on a language level under the notion of a ‘position intentionally held blank’ can be done later.

4 Likes

Not true. I was one of those people, and I am in favour of set literals, but against this proposal.

6 Likes

Fair, I appreciate the correction Paul.

I’m with @pf_moore on this: if Python were a brand new language, {} would be an empty set literal, and likely {:} an empty dict literal. But it’s too late for that. There just isn’t a clean way to spell an empty set literal now short of using Unicode. set() is at least obvious at first glance :wink:.

15 Likes

What is consistency though? Having {/} is consistent in that set() is no longer the only literal container that can’t be written literally whilst empty. But it’s also inconsistent with the other empty literals in that it’s not that same container syntax with the items omitted. And / (or any character) meaning placeholder for nothing if between { and } is consistent with nothing else in Python.

I feel like I say this all the time but consistency for its own sake has no value in itself. It’s valuable when it gives patterns for interoperability or for extrapolating more so that we’re required to learn less. This particular flavour of consistency does nothing in that regard. As long as an empty set is not the obvious extrapolation of a non-empty set without the stuff between the brackets, it’s still a special case to learn. That value does not exist.

11 Likes

It’s also inconsistent with every other literal/display syntax by having to have something between the delimiters in order to declare that it is empty. I’m still strongly against that particular notation, and only weakly in favour of adding any syntax at all.

6 Likes

Given the lack of consensus or convergence to consensus, the community might consider more in depth discussion of the Unicode ∅ symbol.

It’s readable, compact and currently an error so it won’t break anything:

   x = ∅
        ^
SyntaxError: invalid character '∅' (U+2205)

The prior discussion I was able to find was a perfunctory “it’s hard to type.” That’s an editor issue, not a language issue. vim has multiple ways to support it, e.g. a .vimrc setting

inoremap \es ∅

converts \es to the symbol while typing (and it’s fewer keystrokes than set(). I’d expect most other editors to support equivalent functionality.

3 Likes

Noone really wants to use non-ASCII in source code syntax.

<> [] {} () are forwards compatible(?) wrt expanding them beyond the empty instance to being initialised with a population, ie [] becomes [foo, bar]. Somewhat similarly with {/} becomes {foo, bar}.

One can’t do that with a single character like *copies and pastes from your message because I don’t know what the esoteric keyboard combination to produce it is* ∅.

5 Likes

Personally, I’ve just picked up the s = {*()} trick and moved on with life.

10 Likes

I still think the most consensus was around {*} (due to consistency across containers overall; see discussion further up); not sure if @gesslerpd wants to finish the PEP draft he had so that this could be discussed in a separate thread.

1 Like

Consensus among those who chose an option, perhaps. “None of the above” or “Status quo” has far more support.

3 Likes

And those people would remain free to use set(), {*()}, or avoid sets entirely, even if literals got added to address the “hole” that other people would like to see filled one way or another.

If every PEP discussion was aborted by people saying “I don’t need this feature, so no-one should have it”, then nothing would ever change.

Don’t get me wrong; any change needs to stand on it’s own merits, but “I don’t need it” from even a majority is not an argument against it per se.

Exactly.

Correct, but what that means is that consensus on the spelling is not an argument for doing something. And the consensus is quite weak, and primarily shows that there is a lot of disagreement about what would make a good spelling. Consensus among a small subset of people about a spelling is not the same thing as support for the proposal, and there doesn’t seem to be a lot of that.

2 Likes