Hello @ksurma, thanks for the feedback in the PEP 725 thread! I’ll reply here to the items that concern PEP 804. Let me know if you need further clarification and I’ll be happy to elaborate.
[…] adding an indicator that package contains compiled extensions (which would mean adding gcc to the rpm buildroot on our side), would solve the majority of packages for us.
That can be inferred by the presence of compiler definitions in the build table for almost all cases. You could even got one step further and assume that the presence of a build table almost always guarantees that a compiler will be needed. But you don’t need that, since the proposal includes mappings that would convert the necessary definitions into Fedora’s decision to use gcc as the package that provides a C compiler in that distribution.
Mapping: do I understand correctly that if Fedora wants to make use of the PEPs, we have to volunteer to maintain the respective mappings, as they grow?
Yes, that’s what the current proposal suggests: federated maintenance of the mappings. The target communities usually know better (and earlier) about the necessary changes or updates. But even without mappings, the presence of the external tables alone is useful ans it can be the source for better error messaging. For example, upon a setuptools (or any other build backend, really) build error, pip could add a small message saying “this project has an external table with the following dependencies: [list of PURLs]. Their unavailability may be behind the cause of this build error”.
Regarding naming: will there be a semantic hierarchy of mappings (fedora{version}.mapping.json if version is specified, otherwise use fedora.mapping.json?) or are you planning on flat structure that would entail always creating a new versioned file (less preferred by us)? Would we then be able to add more mappings if we need them, e.g. for CentOS Stream, or RHEL?
Yes, see this paragraph:
Each mapping MUST have a canonical URL for online retrieval. Its complete filename MUST be
`{ecosystem-identifier}.mapping.json`, where “ecosystem identifier” MUST conform to this regex:
`[a-z0-9\-_.]+(\+[a-z0-9\-_.]+)?`. In other words, a first field optionally followed by a second,
separated by a + symbol.
The second field is meant to be the version, but a unversion fallback name is recommended. You can add as many mappings as you need, e.g. centos-stream.mapping.json and rhel.json.
looking at dep:virtual/interface/blas in fedora.mapping.json, the key appears 3 times, once for blas, openblas and flexiblas. dep:virtual/interface/lapack can be realized by lapack or flexiblas. Is there any recomendation how tools should interpret multiple entries per id?
For example, adding a priority key to the entry could solve the issue of Fedora defining a preferred/canonical tool for the job.
The current convention followed by the prototype tooling is that the first occurrence wins. The other options are fallback options. But another tool could choose to always prompt the user for interactive workflows. We decided against making this decision on the mapping maintainers side and leave that to the tooling or end user. Having a priority key would mean another field to maintain, keep in sync, and discuss about.
the mappings look arbitrary at some entries (examples mentioned above) and overly specific at others: for example, what are tools supposed to do with a GH repository? Should installers suggest an installation of build dependencies from arbitrary repositories? For rpm builds, this is no-op, as we can only use official rpm packages. How to canonically handle a non-existent mapping?
The Github repositories mentioned in the registry can be considered alternative identities of the same package. Mostly useful for packages that don’t belong in a language-specific registry. e.g. Rust crates in crates.io or JavaScript packages in Npmjs.com are easy to identify because these sites offer a flat namespace. Projects written in C, C++ or Fortran haven’t had such a central location to claim an identifier, so there’s no unambiguous way to refer to them in abstract ways. In those cases, the Github repository URL can serve of additional clarification when a simple generic/<name> is not well known.
It’s unclear to me how the initial registry was compiled, whether it contains only real life use cases or attempts to cover the “whole world”. In the latter case, I’d suggest starting with a minimal set of ids coming from the real life packages, and adding new on demand.
It has been populated by the requirements of real life packages. More specifically, those with external dependencies in the top 150 downloaded packages. More details at GitHub - rgommers/external-deps-build · GitHub.
Technically, the mapping is not specific for the Python land. Should there rather be a overarching registry for all ecosystems dealing with external dependencies? (Not implying that you should be tasked with creating one; just thinking over the best place for the registry).
The text discusses existing options and their unsuitability. We have also discussed the overlap with the PURL maintainers and are monitoring the progress on the prototype purl-registry, but it’s far from ready and is unclear whether it could fit our purposes here in Python land.