Should Python JSON decoders silently convert oversized integers to float?

I came across a behavior in orjson that I’m curious about.

>>> n = 2**64 + 1
>>> s = json.dumps(n)
>>> json.loads(s)
18446744073709551617
>>> orjson.loads(s)
1.8446744073709552e+19

So orjson.loads() turns a JSON integer that doesn’t fit in 64 bits into a Python float.

What’s interesting is the asymmetry with encoding:

>>> orjson.dumps(2**64)
TypeError: Integer exceeds 64-bit range

I understand that orjson uses yyjson, which has a 64-bit integer representation and falls back to double on integer overflow.

But should a Python-facing JSON decoder silently change the type from int to float, given that Python int is arbitrary precision? Would rejecting the value or preserving it as an int be more expected?

I’m mainly interested in the general Python/API-design question rather than whether orjson specifically should change.

If it’s possible to provide the precision, it should be done. JSON (AFAIK) still doesn’t define the number type/limits, so as long as it’s not a DoS concern it should be integers if no fractional part is present.

This is the kind of behaviour that would cause bugs for me, because I’d never expect this to be happening. (And I don’t read the docs :wink:)

With ‘hindsight’ I can understand where it comes from. A lot of people view Jsons as representing javascript objects, rather than just being strings. So they view the ‘proper’ way to read json to first convert to a javascript-like types, and then to Python.[1]


  1. I view jsons as strings, therefore capable of representing integers of arbitrary size, and natively representing Decimal numbers (and floats only as a by-product). ↩︎

Correct, it’s one of the most remarkably underspecified popular standards there is. Somehow it doesn’t cause massive interoperability problems :slight_smile:

Personally I think it’s better to faithfully reflect the data, even if JavaScript wouldn’t have been able to; though with a pragmatic limit (I don’t, for example, think that we should emit decimal.Decimal for floats by default, but I’m glad there’s an easy way to do it if you want it). JSON does not strictly represent JavaScript data, and it should be possible to carry large integers through it. If I encode a large integer in Python, I can decode it in Pike, even though JS wouldn’t be able to handle it.

Whether a third-party orjson package follows this principle is up to its maintainers, though. Maybe its purpose is performance rather than maximum generality. That’s a valid tradeoff and a reason for having different JSON packages. But my default expectation is that JSON can carry large integers without loss of precision.

Conversion to float seems like a bad idea. It makes the encode/decode step lossy and irreversible. It may also be a schema violation if someone is using a typed dict.