# Is the size comparison of utf-8 and utf-16 valid?

**URL:** <https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614>\
**Category:** Core Development\
**Tags:** help\
**Created:** [September 11, 2024, 2:50pm UTC](https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614 "2024-09-11T14:50:38Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![byundojin](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/byundojin/32/22483_2.png) [@byundojin](https://discuss.python.org/u/byundojin)\
**Post date:** [September 11, 2024, 2:50pm UTC](https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614/1 "2024-09-11T14:50:38Z")

</div>

If function \_io\_\_WindowsConsoleIO\_write\_impl does not produce as expected output from WriteConsoleW, it is the process of recalculating the existing len.

At this point, function WideCharToMultiByte finds the byte length of utf-8 and function MultiByteToWideChar finds the letter length of utf-16. And is this comparison valid for comparing the results of these two?

> <https://github.com/python/cpython/blob/6e23c89fcdd02b08fa6e9fa70d6e90763ddfc327/Modules/_io/winconsoleio.c#L1049-L1062>

---

<div class="post-metadata">

**Author:** ![steve.dower](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/steve.dower/32/56_2.png) [@steve.dower](https://discuss.python.org/u/steve.dower)\
**Post date:** [September 11, 2024, 3:25pm UTC](https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614/2 "2024-09-11T15:25:34Z")

</div>

This is apparently code that I wrote (many years ago), but I don’t recall it, and the whole quoted block seems like a weird thing to be doing.

Basically, for those who haven’t clicked through to read the whole function, `write(b)` is only allowed a single system call, and if that call doesn’t write the entire buffer it has to return the number of bytes that were actually written. `winconsoleio` converts back into Unicode to write it out. At this point, the `n < wlen` means that fewer (UTF-16) characters were written than were in the buffer, and so we should be recalculating the number of bytes that were actually used to return (in `len`).

An assert that would probably make sense here would be to check the new `wlen` against `n`, to verify that the amount of data we’re saying was used is able to round-trip through the UTF-8-\>UTF-16 conversion. Comparing to `len` is correct if the buffer only contains single-byte characters, which we have no way of knowing here.

But checking the length a second time doesn’t really seem necessary. Updating `len` with the `WideCharToMultiByte` result here should be enough I think. Lines 1057-1061 aren’t necessary.

---

<div class="post-metadata">

**Author:** ![byundojin](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/byundojin/32/22483_2.png) [@byundojin](https://discuss.python.org/u/byundojin)\
**Post date:** [September 11, 2024, 11:58pm UTC](https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614/3 "2024-09-11T23:58:04Z")

</div>

I’ll create a PR right away then

---

<div class="post-metadata">

**Author:** ![storchaka](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/storchaka/32/217_2.png) [@storchaka](https://discuss.python.org/u/storchaka)\
**Post date:** [September 12, 2024, 4:19am UTC](https://discuss.python.org/t/is-the-size-comparison-of-utf-8-and-utf-16-valid/63614/4 "2024-09-12T04:19:27Z")

</div>

It looks to me that the intented assert is `wlen == n`.
