PEP 842: Module Exports (new revision)

Thanks for the update! Here are a few nits.

  • You fixed the example to raise ExportError instead of issue a warning. But you still state in the next sentence “This is **not** intended to be an access modifier for Python; see the rationale .” Much later you explain what you mean by that – there are legitimate ways to bypass the export controls. But does “access control” really mean to make it physically impossible to access something? That would be hard in Python, as long as we have the ctypes module (or we can load unvetted extensions). Maybe we should simply not use the term “access control” and just state up front what we mean – that there are ways around an ExportError?

  • Do we really need to add words to explain that you can’t say def export name... even though the grammar makes it abundantly clear?

  • Maybe we should use an example to explain why reusing __all__ isn’t backwards compatible.

  • The explanation of how we went from ImportError to ExportWarning and then back to ExportError should clarify more why we went full circle. Or maybe instead of describing the back-and-forth, explain how issuing a warning is worse due to the inherently poor ergonomics of warnings – warnings occurring in a library punish the user for something they can’t fix, and testing frameworks like pytest turn warnings into errors.

  • I’d shorten the section on private to explaining why this is out of scope: it’s a different set of requirements that deserve a separate discussion and another PEP.

  • In the section about access to internals from within the same package, consider renaming _utils.py to utils.py – the preferred name if we had a proper solution for this. It’s fine that you don’t mention my proposal yet – it reeks expensive, and although a C implementation could be much faster, it would still be a lot more expensive than currently. We might have to benchmark this before accepting it.

  • I think another open issue is whether we need a marker near the top of a file indicating that export is being used, and if so, what that marker should look like, and what should happen when the marker is absent but export is used.

1 Like

I don’t think it matters too much if it’s an Error or a Warning, both make declaring the API with the new features a breaking change[1].

I generally agree with the principle that is even mentioned in the PEP that access to internals is a feature when you need it. So from my perspective this PEP is trying to remove what I find to be a useful feature[2].

This doesn’t mean that we can’t try to steer people away from using internals and provide ways to better declare what the internals are, but I think it’s better to put up a sign than a roadblock.

An indicator for linters and language servers that they should not show these names and associated change to dir / help both doesn’t require breaking currently functioning code and also correctly targets the developer who’s accessing the internals, rather than anyone further downstream who can’t do anything about it.


I think that the examples given are a good case for why not to provide this statement and not to make it an error.

First, the examples given all appear to be from the case where a module has been unintentionally re-exported. Do you have examples of someone importing an internal feature of a library that isn’t a re-export where this has caused a problem (intentionally or unintentionally)?

Second, the result of libraries making this change and removing the unintentional export is that end users then need to find workarounds for slow updating dependents that happened to depend on those unintentional exports. You see this in the scikit-learn examples where people give sys.modules["sklearn.externals.six"] = six as the workaround. Enforcing the ExportError is likely going to end up creating the need for new, more convoluted workarounds in these situations.

The result is that using the new feature with the error in place means projects have to choose between either intentionally breaking downstream projects and their users or not declaring their API with the new public export modifier, which feels like a bad choice either way.


  1. In the practical sense that you’re breaking libraries that currently use un-exported internals and anything which depends on these libraries that may not be able to do anything about it other than complain and potentially pin the old, unsupported version of your library. ↩︎

  2. I’m not going to entertain the argument that this isn’t an access modifier. By raising an exception on the current standard way of importing something that is internal, it acts as an access modifier in practice. ↩︎

I think that’s a good idea. That sentence in the abstract is mostly there to prevent the knee-jerk reaction of “Python doesn’t have private variables!”

I added that in response to this comment.

Other comments acknowledged; I’ll update the PEP.

Yes, here are some:

logging._acquireLock / logging._releaseLock

This one was partially our fault. The logging docs had an example using these names, which many users mindlessly copied and used in their own projects, without any idea that they were touching internals. They were broken in Python 3.13 when the name was removed:

logging._levelNames

The only way to turn something like "DEBUG" into 10 was to use this. This broke things in 3.4 when this was split into _nameToLevel/_levelToName:

asyncio.coroutines._is_coroutine

Setting func._is_coroutine = asyncio.coroutines._is_coroutine was the only way to make a sync function returning an awaitable pass iscoroutinefunction. Copy-pasting between projects made this pattern common.

pathlib._Accessor / pathlib._Flavour

In older versions, the only way to subclass a Path object was to use the internal API. This led to a lot of breakage when they were removed in favor of native Path subclassing:

matplotlib.cbook._check_in_list / matplotlib.cbook._rename_parameter

matplotlib kept internal helper functions in a public module. People used them anyway, and they were broken when they were removed:

concurrent.futures.thread._threads_queues

In Python 3.8, many people wanted to make CTRL+C kill a ThreadPoolExecutor. The recipe that was shared around used the internal API, which broke in 3.9 when worker threads stopped being daemon.

sysconfig._get_default_scheme

This was a way to determine where Python installed things. It was found to be useful, so it was renamed to get_default_scheme in Python 3.10. This broke things that had already been relying on the old name:

re._pattern_type

The easiest way to get the type of a compiled regular expression (i.e., not just type(re.compile(''))) was to use this, which broke code in 3.7 when it was removed:

It’s trying to make it much harder to do it accidentally, and to hopefully encourage people to ask for public APIs upstream instead of just using something private and relying on it. As I’ve said many times, this PEP does not prevent you from accessing private names. If you decide you don’t care, then just del module.__export__ (or module.__export__.remove("name_you_want").

No, it just changes when the workarounds have to occur.

You’re envisioning a case where a library has private names, downstream users begin relying on those private names, and then the library uses this PEP to make private names raise an ExportError on access. Downstream users who were already relying on internals would be broken by this, but arguably, that was going to happen anyway when the library inevitably broke the private API. It’s a one-time break, though; once users have migrated (or pinned), they are strongly discouraged from ever making the mistake again.

3 Likes

From the discussion thus far, I get the impression that we can split people who use private APIs into three main groups:

  1. People who do this truly accidentally—e.g., because their IDE’s autocomplete suggests a private API that sounds like it might do what they’re looking for.
  2. People who deliberately use a private API and are aware of the risks. (This includes internal use within a package.)
  3. People who are copy-pasting a code snippet (usually from someone in group 2) to fulfil a specific need, without thinking too much about the code snippet.

For group 1, I think this PEP would indeed be helpful once it’s supported by IDEs.

For group 2, this PEP would not change much; they will continue to do this, e.g., using the del module.__export__ workaround. (Adding a single line of boilerplate to get the expected result right now is still orders of magnitude less friction than posting an issue upstream to ask for a public API and then waiting months to years for this to be accepted, implemented and released; so I don’t think this will be sufficient encouragement to make a practical difference.)

For group 3, I also don’t think that this PEP would change much. The code snippet they copy is now going to be one line longer; but that one del module.__export__ line does not look “scary” enough to be an effective deterrent.[1]

Framed like this, the question to me becomes whether group 1 actually makes up a large fraction of private API usage. If they do, then this PEP could make a significant difference and I would support it. If, however, groups 2 and 3 dominate, then I don’t expect this PEP to make a significant difference and, to me, the small improvement it offers w/r/t group 1 is outweighed by the additional boilerplate and churn it introduces.


  1. And I suspect there is little appetite to rename __export__ to something sufficiently scary-sounding, jax-style :wink: ↩︎

3 Likes

This will change the semantics of the module globally and is not really appropriate for a library to do because other libraries in the application, or the application itself, may want those warnings for exceptions. I’ve debugged a few issues in production that were hidden by an “offending” library disabling a global warning another library emitted (sqlalchemy warning) and not having those warnings led to incorrectly eliminating possible causes (because isolated tests without the “offending” library emitted the warning leading to thinking that wasn’t the cause. I am vehemently opposed to libraries making global changes that effect how other libraries may behave.

7 Likes

I figured that, but I don’t think that question warrants a mention in the PEP. The answer is just “it’s already in the PEP, see ”. Since the PEP has the answer already, it doesn’t have to repeat itself.

But, if you want to keep the section, at least use the example from the post (async export def...). There are no examples of modifiers going after def.

1 Like

Thanks for the examples. At least some of these should probably be in the PEP.

Many of these also fall into the case of people needing to use these internals because we didn’t provide a public API. I don’t think we remove the need to do these things by making it more difficult to do so.

The sysconfig._get_default_scheme example would work as an argument for a way to mark something that isn’t prefixed with an underscore as private - but not to prevent accessing it. The feature was useful, the breakage was in renaming.

Sometimes you do ask for public APIs upstream[1] but you’re not the maintainer and you can’t just add the API yourself so you need a way to get what you need done.

From my perspective it’s not that I “don’t care” it’s that the only option to implement the feature I need is to use the internals. Sure I’d love if there was a public method to do so, but there isn’t and getting a public API soon didn’t seem likely.

I don’t really think it’s good practice for a library to mutate global state in other modules. For users yes we’ll probably find recommendations to remove the __export__ name in order to make things work again.

Your second method appears to be the opposite, wouldn’t that make “name_you_want” private?

I guess there would be a disagreement on inevitability. By creating this export feature you’re pretty much forcing that and I don’t think this is necessary or even good.

You give a lot of examples where the only good choice was to make the mistake that you’re now saying users should never make again.


  1. Note that I don’t think the internals I need to implement this need to be public, and as such I wouldn’t propose making them so. But I do need to access those internals. ↩︎

4 Likes

sys._getframe() is another example of a name underscored not to signal privateness, but instead that “you’re doing something special, non-standard, and potentially non-portable.”

2 Likes

Yes, I’m aware – you asked for some examples where the usage was intentional. I was unsure whether you cared about the examples themselves or whether they applied to the PEP.

I know. I don’t think mutating __export__ should be considered good practice. I think it should be a code smell that prompts the reader to ask why it’s there, and if there’s a better solution. (This applies to @Jost’s comment from above too; I’d like to believe that library developers won’t blindly del module.__export__.)

The mistake is that they’re (seemingly) opting into breakage without realizing it. If you genuinely need a private API, then you should either ask upstream or pin/vendor the module, and this PEP makes that process more explicit. There shouldn’t be cases where people use private APIs and then continue to update without expecting breakage.

I don’t think we recommend that libraries pin or vendor dependencies due to the other issues it can create. I know an upstream change might break something - to some extent that’s always the case anyway, just more likely with internals.

2 Likes

So what is the workaround for library developers who need to use internals of a dependency? Pinning to the old version isn’t realistic - no-one is going to be happy potentially blocking security updates. And asking for a public API is also not realistic (in the short tern, at least) - most projects are volunteer-maintained, and you could be waiting a long time for such a request to be implemented.

1 Like

I just read the PEP and I’m confused about what exception would be raised by

from foo import private

Would it be ExportError or ImportError?

The PEP says that

import foo
foo.private

would raise ExportError which is a subclass of AttributeError.

I assume that the implementation of from foo import nonexistent internally turns an AttributeError into an ImportError. So is the idea that from foo import private would turn the ExportError into an ImportError?

I’m thinking of code that looks like

try:
    from foo import new
except ImportError:
    # foo < 1.0
    def new_fallback(...) ...

Would that sort of thing need to catch both ImportError and ExportError or would it still just be ImportError?

1 Like

I agree it should not be considered good practice. Just like it shouldn’t be considered good practice to use undocumented private APIs. However, the whole reason we’re having this discussion is because some small fraction of users still do these things.

So, yes, del module.__export__ is a code smell. But is it a sufficiently alarming code smell to people in group 3 (who, by definition, are undeterred by other code smells like use of module._private_api()) that it would deter a large fraction of them? I don’t think it is.

More generally: If we want to have a workaround that does not place an undue burden upon group 2, this workaround would also be easy enough for group 3 to copy & paste, so it would not serve as an effective deterrent. On the other hand, if we do not have an easy workaround, there will be significant pushback from group 2 (which is probably overrepresented here on DPO).

1 Like

You don’t have to always vendor the entire dependency. In a lot of cases, just copy-pasting a certain class or function from the source is enough, and that’s not prone to breakage across updates.

Fundamentally, I don’t think this should turn into a discussion about “what if I need internals?”, because that is not the problem this PEP is trying to solve, or the solution that it’s trying to remove. There are clear workarounds if you want to continue opting in to breakage.

I care most about people accidentally using internal APIs from the standard library and having to deal with the breakage that ensues when we change things. (If this PEP is withdrawn, I may pursue the idea for the stdlib on its own, because that’s where unprotected access to internals has reliably proven to be a problem in practice.)

You use __dict__ if you don’t want to change global state. But I don’t buy the “blocking security updates” argument – that argument might make sense for pip, but for many libraries, the chance of breakage by using internals is much higher than security vulnerabilities, especially if you’re not using a given library on anything that could be a potential attack surface. (Are you really that concerned about a security issue in sysconfig, for example?)

I can just make it inherit from both ImportError and AttributeError.

I wrote a section in the PEP about why module._private_api isn’t a universal indicator that something is private. In short, del module.__export__ is absolutely crystal clear that you’re using something internal; an underscored name is not.

1 Like

I don’t see why you combine internal use with downstream use here. Those are completely different cases. The primary reason to make something private is to send a clear signal to downstream that internal changes in the package can break use of the private thing.

This makes no sense for internal use within a package. Why would a package author define __export__ in one module only to put del module.__export__ in another module? If they delete __export__ then downstream code won’t see the ExportError either so why define __export__ at all?

1 Like

I adapted the “semantic implementation” section of the PEP and it does change the ExportError into ImportError:

# mod.py

public = 1
private = 2

__export__ = ['public']

import sys

class ExportError(AttributeError):
    pass

if "__all__" not in globals():
   __all__ = __export__

def _is_dunder_name(name):
    return (len(name) > 4) and name.startswith("__") and name.endswith("__")

# Attributes not in the __dict__ fall back to the normal lookup
class Mod:
    def __getattribute__(self, name):
        try:
            value = globals()[name]
        except KeyError:
            raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None

        if _is_dunder_name(name):
            return value

        if name not in __export__:
            raise ExportError(f"{name!r} is not exported by {__name__!r}")

        return value

sys.modules[__name__] = Mod() # type: ignore

def __dir__():
   names = []
   for name in globals().keys():
      if (name in __export__) or _is_dunder_name(name):
         names.append(name)
   return names

Then

>>> import mod
>>> mod.public
1
>>> mod.private
...
mod.ExportError: 'private' is not exported by 'mod'
>>> from mod import private
...
ImportError: cannot import name 'private' from 'mod' (.../mod.py)

Either way the PEP should say explicitly what exception from foo import private would raise and what should be the expected way to catch missing imports if __export__ is potentially being used.

1 Like

I’d like to note that I do think this is a situation that’s worth improving but I don’t think the syntax and the error help. If there is a public API a developer should be using ideally we’d want their tooling to point them towards the correct API (assuming there is one), rather than just denying the private API that they may have heard about.

2 Likes

I have a question about what is and isn’t allowed by the export assignment syntax. It gives examples of assigning to a single name or multiple names, and says that augmented augmented are disallowed. It doesn’t say anything about subscripts (x[0] = something) or attributes (x.y = something). Are those allowed to be used with export, and if they are, what would it mean? What about grouping names with () or []? Is that permitted? (I think the syntax is ambiguous if [] is allowed.)

What would be your preferred approach?

No, they’re not supposed to be allowed. I’ll clarify this in the PEP – thanks for the question!

1 Like

This is a curiosity driven discussion relating to the three current PEPs 842, 843 and 844. I have tried to be fact focussed, and to put my personal opinions to one side. Mostly, this post is a sequence of relevant facts. I have tried to use neutral and fair language.

The Python docs says

The public names defined by a module are determined by checking the module’s namespace for a variable named __all__; if defined, it must be a sequence of strings which are names defined or imported by that module.

The abstract of PEP 842 says

This PEP proposes an export statement that modules can use to express intent about the visibility of variables from outside the module.

What is meant by public name? I don’t think the Python docs answer this question. According to RealPython

In Python, a public name refers to any attribute, method, variable, class, object, or module that is intended to be used from outside its immediate scope.

Python doesn’t have a strict mechanism to distinguish between public and private names like Java or C++ have. Therefore, it uses the terms public and non-public names.

If this is the correct meaning for public name, then perhaps __all__ are the top level names in a module that are intended for use outside the module. And export is a means of expressing intent about visibility outside the module.

So now I ask, what’s the difference between:

  • intended for use outside the module (__all__)
  • Intended to be visible outside the module (export)

I don’t have an answer for this, but I have an example that involve both concepts.

Consider asyncio/_init_.py. Briefly, it goes like this:

  1. Import the sys module.
  2. Import * from submodules of asyncio.
  3. Define __all__ to be the sum of the __all__ in the submodules.
  4. Do an if ... then ... based on the platform. This adds entries to the module __all__.
  5. Give the module a custom __getattr__ related to loop policies.

In short, the names in the module are:

  1. The names in module’s __all__.
  2. The names that are in even the so-to-speak empty module.
  3. The name __getattr__ .
  4. The name sys.

This is pretty good, but I see three problems with the definition of __all__ in asyncio.__init__.py. The problems are:

  1. Someone who reads the source file does not get an explicit list of names.
  2. Summing the submodule __all__ fails if they have different types. You can’t add a list to a tuple or, for what it’s worth, a set.
  3. There are names in __all__ that surely don’t belong there.

Regarding (3) see below, which is clearly related to (1):

>>> for name in sorted(asyncio.__all__):
...     if name[0] == '_' and name[1] != '_':
...         print(name)
...         
_enter_task
_get_running_loop
_leave_task
_register_task
_set_running_loop
_unregister_task
>>> 

Each of these is a builtin function, and so probably from a C-coded module.

>>> set(type(getattr(asyncio, name)) for name in aa)
{<class 'builtin_function_or_method'>}