# Archspec: a library for labeling optimized binaries

**URL:** <https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149>\
**Category:** Packaging\
**Created:** [February 9, 2020, 6:51am UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149 "2020-02-09T06:51:26Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tgamblin](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/tgamblin/32/1211_2.png) [@tgamblin](https://discuss.python.org/u/tgamblin)\
**Post date:** [February 9, 2020, 6:51am UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/1 "2020-02-09T06:51:26Z")

</div>

There hasn’t (so far) been a standard for naming and comparing optimized binaries; most systems just call things by the ISA family (e.g. `x86_64` or `ppc64le`), without much more information.

We’ve pulled a library out of [Spack](https://github.com/spack/spack) that could be used to label optimized wheels. We’re calling it archspec:

> **[archspec/archspec](https://github.com/archspec/archspec)**
>
> A library for detecting, labeling, and reasoning about microarchitectures - archspec/archspec

The library does several things that will be of interest to packagers:

1. Detects the microarchitecture of your machine (i.e., not just `x86_64`, but `haswell`, `skylake`, `thunderx2`, or `power9le`.
2. Compares microarchitectures for compatibility. You can say things like `skylake > haswell` to test whether a `skylake` machine can run `haswell`-optimized binaries (you’ll get `True`).
3. Query the features available on a particular microarchitecture. You can ask, e.g. `'sse3' in haswell` or `'neon' in thunderx2`
4. Ask what compiler flags to use to get different compilers to output binaries for a specific target. e.g., you can ask what flags to use on `gcc`, at a particular version, to get a binary for `haswell`.

The library also defines a set of canonical and hopefully familiar names for microarchitectures. The list currently looks like this (from spack output):

```auto
$ spack arch --known-targets
Generic architectures (families)
    aarch64 arm ppc ppc64 ppc64le ppcle sparc sparc64 x86 x86_64

GenuineIntel - x86
    i686 pentium2 pentium3 pentium4 prescott

GenuineIntel - x86_64
    nocona nehalem sandybridge haswell skylake skylake_avx512 cascadelake
    core2 westmere ivybridge broadwell mic_knl cannonlake icelake

AuthenticAMD - x86_64
    k10 bulldozer zen piledriver zen2 steamroller excavator

IBM - ppc64
    power7 power8 power9

IBM - ppc64le
    power8le power9le

Cavium - aarch64
    thunderx2

Fujitsu - aarch64
    a64fx

```

We use this library in Spack to ensure:

1. that every binary package is built for a specific target
2. that we can use a particular binary package on a given host machine.

The detection logic and library bindings are currently for Python, but the domain knowledge (features, arch names, etc.) is all in a generic [json file](https://github.com/archspec/archspec-json) with a schema, which we hope can enable people to easily build other language bindings. Currently the library supports macOS and Linux (Windows help would be great).

We’d love to get more contributions to keep the data up to date, and we’re hoping that if this takes off, it’ll enable people to more easily distribute optimized binary packages and containers.

Comments/suggestions/contributions welcome! For more info, see [this talk from FOSDEM](https://fosdem.org/2020/schedule/event/archspec/).

---

<div class="post-metadata">

**Author:** ![sumanah](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/sumanah/32/1722_2.png) [@sumanah](https://discuss.python.org/u/sumanah)\
**Post date:** [February 10, 2020, 11:38pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/2 "2020-02-10T23:38:07Z")

</div>

This is SO GREAT (in my opinion)!!

---

<div class="post-metadata">

**Author:** ![arcivanov](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/arcivanov/32/15454_2.png) [@arcivanov](https://discuss.python.org/u/arcivanov)\
**Post date:** [February 14, 2020, 5:06pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/3 "2020-02-14T17:06:15Z")

</div>

There is an issue with limiting itself to the architecture due to the fact that large cloud providers order themselves customized CPUs with some of the features missing from the architecture. I had numerous cases even going back to 2010 where you’re compiling for AWS architecture and get an Illegal Instruction fault, dig in into the spec and find that a specific instruction set is missing from the otherwise standard architecture.

So you technically need to introduce an arch-specific CPU capability bit field and do an `and` with available distros’ bitfields to see which ones are maximally compatible and do a fallback.

---

<div class="post-metadata">

**Author:** ![tgamblin](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/tgamblin/32/1211_2.png) [@tgamblin](https://discuss.python.org/u/tgamblin)\
**Post date:** [February 14, 2020, 5:22pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/4 "2020-02-14T17:22:51Z")

</div>

The library determines the architecture by the available features, so if a cloud provider disables key features (or at least the ones that archspec models), then we detect the platform as a different architecture – e.g., if you were to disable avx-512 features on `skylake`, we’d identify the machine as a `haswell` and only use `haswell` binaries. So from that perspective, the features you’re talking about are already modeled.

---

<div class="post-metadata">

**Author:** ![arcivanov](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/arcivanov/32/15454_2.png) [@arcivanov](https://discuss.python.org/u/arcivanov)\
**Post date:** [February 14, 2020, 10:34pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/5 "2020-02-14T22:34:08Z")

</div>

Then how do you designate architecture that is used in the cloud-specific CPUs?

---

<div class="post-metadata">

**Author:** ![arcivanov](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/arcivanov/32/15454_2.png) [@arcivanov](https://discuss.python.org/u/arcivanov)\
**Post date:** [February 14, 2020, 10:34pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/6 "2020-02-14T22:34:47Z")

</div>

Because optimizing for `haswell` is not the same as optimizing for `skylake-avx512`.

---

<div class="post-metadata">

**Author:** ![arcivanov](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/arcivanov/32/15454_2.png) [@arcivanov](https://discuss.python.org/u/arcivanov)\
**Post date:** [February 14, 2020, 10:38pm UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/7 "2020-02-14T22:38:05Z")

</div>

There is also PyTorch cpuinfo library that provides a very precise CPU detection feature, including SoC levels.

> **[pytorch/cpuinfo](https://github.com/pytorch/cpuinfo)**
>
> CPU INFOrmation library (x86/x86-64/ARM/ARM64, Linux/Windows/Android/macOS/iOS) - pytorch/cpuinfo

---

<div class="post-metadata">

**Author:** ![tgamblin](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/tgamblin/32/1211_2.png) [@tgamblin](https://discuss.python.org/u/tgamblin)\
**Post date:** [March 1, 2020, 12:19am UTC](https://discuss.python.org/t/archspec-a-library-for-labeling-optimized-binaries/3149/8 "2020-03-01T00:19:51Z")

</div>

> [@arcivanov](#):
>
> Then how do you designate architecture that is used in the cloud-specific CPUs?

If there are features missing from the architecture, it needs its own name in this model. So we would add a cloud-specific CPU name (e.g., `graviton2`).

If the chip _doesn’t_ have the features we expect, we’ll build for the next best thing we can find (see [Incorrect arch detection · Issue #15151 · spack/spack · GitHub](https://github.com/spack/spack/issues/15151)), so if the cloud chip isn’t in our list, we’ll pick something else.

The goal here is to have optimized binaries _and_ to know where we can reuse them. Not necessarily to build as natively as possible for the host. We’re sacrificing some degree of specificity by having well defined names. But like I said, we should add the cloud-specific CPUs. We don’t have to only support names that, e.g., Intel defines.
