# How to convert character to hexadecimal and back?

**URL:** <https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056>\
**Category:** Python Help\
**Created:** [March 22, 2023, 9:55am UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056 "2023-03-22T09:55:56Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pejamide](https://avatars.discourse-cdn.com/v4/letter/p/f07891/32.png) [@Pejamide](https://discuss.python.org/u/Pejamide)\
**Post date:** [March 22, 2023, 9:55am UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/1 "2023-03-22T09:55:56Z")

</div>

I tried to code to convert string to hexdecimal and back for controle as follows:

```auto
input = 'Україна'
table = []
for item in input:
  table.append(item.encode('utf-8').hex())
output = ''
for item in table:
  output += chr(int(item, 16))
print(output)

```

To my surprise the output is total different from the input. What have I done anything wrong?

---

<div class="post-metadata">

**Author:** ![Pejamide](https://avatars.discourse-cdn.com/v4/letter/p/f07891/32.png) [@Pejamide](https://discuss.python.org/u/Pejamide)\
**Post date:** [March 22, 2023, 9:57am UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/2 "2023-03-22T09:57:11Z")

</div>

The output was: 킣킺톀킰톗킽킰

---

<div class="post-metadata">

**Author:** ![Rosuav](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/rosuav/32/3429_2.png) [@Rosuav](https://discuss.python.org/u/Rosuav)\
**Post date:** [March 22, 2023, 10:16am UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/3 "2023-03-22T10:16:44Z")

</div>

> [@Pejamide](#):
>
> To my surprise the output is total different from the input. What have I done anything wrong?

UTF-8 is one particular encoding. If you want a reversible transformation, the chr function’s counterpart is ord.

---

<div class="post-metadata">

**Author:** ![cameron](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/cameron/32/2658_2.png) [@cameron](https://discuss.python.org/u/cameron)\
**Post date:** [March 22, 2023, 10:40am UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/4 "2023-03-22T10:40:25Z")

</div>

> I tried to code to convert string to hexdecimal and back for controle  
> as follows:
> 
> ```auto
> input = 'Україна'
> table = []
> for item in input:
> table.append(item.encode('utf-8').hex())
> 
> ```

At this point you have a list of hexadecimal strings, one per character  
(well, technically Unicode codepoints). Example:

```
 >>> input = 'Україна'
 >>> hexcodes = [item.encode('utf-8').hex() for item in input]
 >>> hexcodes
 ['d0a3', 'd0ba', 'd180', 'd0b0', 'd197', 'd0bd', 'd0b0']

```

However, utf-8 is a variable width multibyte encoding. Its value is that  
for the first 128 codes (the ASCII range) the byte encoding is the same  
at 1 byte per code. (This made plain ASCII files automatcally UTF-8  
compatible and made a lot of western european text compactly  
represented. The flip side is that later values have a longer encoding.)

The reverse of this encoding is not reversing your `.hex()` call. The  
high order bits indicate the length of the encoding of the code value,  
and do not themselves contribute to the code value itself.

To undo this:

- undo your `hex()` into bytes
- decode the bytes by decoding as UTF-8: `bs.decode('utf-8')` if your  
bytes values was in a variable named `bs`

Cheers,  
Cameron Simpson [cs@cskk.id.au](mailto:cs@cskk.id.au)

---

<div class="post-metadata">

**Author:** ![franklinvp](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/franklinvp/32/11142_2.png) [@franklinvp](https://discuss.python.org/u/franklinvp)\
**Post date:** [March 22, 2023, 12:27pm UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/5 "2023-03-22T12:27:03Z")

</div>

> [@Pejamide](#):
>
> ```auto
> output = ''
> for item in table:
> output += chr(int(item, 16))
> 
> ```

This is a separate suggestion.

In Python a `str`, like the value of `output` is immutable. So, when we do `output += chr(int(item, 16))` a new `str` gets created each time and destroyed after the next iteration of the `for` loop. We only really need the final string to be created.

You could use `str.join` like

```auto
output = ''.join(chr(int(item, 16)) for item in table)

```

The use of [list comprehension](https://docs.python.org/3/tutorial/datastructures.html#list-comprehensions) can also be applied to the creation of `table`, as

```auto
table = [item.encode('utf-8').hex() for item in input]

```

---

<div class="post-metadata">

**Author:** ![kknechtel](https://avatars.discourse-cdn.com/v4/letter/k/e47c2d/32.png) [@kknechtel](https://discuss.python.org/u/kknechtel)\
**Post date:** [March 22, 2023, 8:18pm UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/6 "2023-03-22T20:18:06Z")

</div>

Code style hints, first:

1. Please do not use `input` as a variable name - that causes [shadowing](https://stackoverflow.com/questions/20125172), meaning that the built-in function `input` is no longer available (the name `input` can only mean one thing at a time).

2. As Franklin suggested, consider using list comprehensions, generator expressions etc. to iterate and collect data - it’s much simpler and more direct. The code could be as simple as:

```auto
data = 'Україна'
table = [item.encode('utf-8').hex() for item in data]
output = ''.join(chr(int(item, 16)) for item in table]
print(output)

```

> [@Pejamide](#):
>
> ```auto
> for item in input:
> table.append(item.encode('utf-8').hex())
> 
> ```

This means that each character in the input string will be converted into bytes **using the UTF-8 encoding** , and then a string representing those byte values will be created. Each such value is added to the list.

> [@Pejamide](#):
>
> ```auto
> for item in table:
> output += chr(int(item, 16))
> 
> ```

This means that each of the hex strings will be converted into a single integer, and then the corresponding Unicode code point will be looked up.

The reason this does not give the same result is because _[UTF-8 encoding](https://en.wikipedia.org/wiki/UTF-8) does not convert characters into the bytes used for an integer representation of that element of the string_. It uses a _variable amount of bytes_ for each element, and sets some “flag” bits as a way of signalling, in-band, how many bytes to use.

For example, `'У'` contains a single element with Unicode code point 1059. Stored as a 2-byte integer, that would require the bytes 0x23 0x04 in little-endian, or 0x04 0x23 in big-endian. UTF-8 is conceptually “big-endian”, but it **also** sets some flag bits, in such a way that the encoding is instead 0xd0 0xa3 - as an integer, 53411.

Two-byte UTF-8 sequences use **eleven** bits as actual information-carrying bits. Three bits are set in the first byte to mean “this is the first byte of a 2-byte UTF-8 sequence”, and two more in the second byte to mean “this is part of a multi-byte UTF-8 sequence (not the first byte)”. (This is a bit redundant, but encoding this way means that it’s easy to detect corruption when a code point gets sliced in half).

To undo the encoding, we should instead get the corresponding **bytes** from the hex dump (rather than a single integer), and **decode** it:

```auto
input = 'Україна'
table = []
for item in input:
  table.append(item.encode('utf-8').hex())
output = ''
for item in table:
  output += bytes.fromhex(item).decode('utf-8')
print(output)

```

To get hex dumps of the actual Unicode code point values, we should use `ord` as the opposite of `chr` (which gives an integer rather than `bytes`), and convert the integer to a hex dump using string formatting:

```auto
input = 'Україна'
table = []
for item in input:
  table.append(f'{ord(item):x}')
output = ''
for item in table:
  output += chr(int(item, 16))
print(output)

```

Note that this approach will use whatever number of hexadecimal digits is needed to represent the values (even, as here, if that’s an odd number).

---

<div class="post-metadata">

**Author:** ![Pejamide](https://avatars.discourse-cdn.com/v4/letter/p/f07891/32.png) [@Pejamide](https://discuss.python.org/u/Pejamide)\
**Post date:** [March 22, 2023, 8:35pm UTC](https://discuss.python.org/t/how-to-convert-character-to-hexadecimal-and-back/25056/7 "2023-03-22T20:35:06Z")

</div>

Please give me your code as an example how I can reverse it, as I understand that code much better.  
Thank you very much, as I’m not expert.
