To save a list of strings as bytes and get exactly the same list back, you can't just join them with a separator: the strings might contain the separator. The standard fix is to write each string's length first, so the reader knows how many bytes to take.
Implement two functions:
encode(strings: list[str]) -> bytesdecode(data: bytes) -> list[str]
such that decode(encode(xs)) == xs for every list of strings, including empty strings, newlines, commas, NUL characters and non-ASCII text.
Use this exact format so other programs could read it. For each string, in order:
- Encode the string as UTF-8.
- Write the number of bytes (not characters) as a 4-byte big-endian unsigned integer:
struct.pack(">I", n). - Write the UTF-8 bytes.
Nothing else: no header and no count of strings.
encode(["hi", ""]) # b"\x00\x00\x00\x02hi" + b"\x00\x00\x00\x00"
encode(["é"]) # b"\x00\x00\x00\x02\xc3\xa9" ("é" is 1 character but 2 bytes)
decode(encode(["a,b", "line\nbreak"])) # ["a,b", "line\nbreak"]
decode(b"") # []
If data ends in the middle of a length or a string (a truncated file), decode must raise ValueError rather than return a partial result.
Show hint
decode walks an offset through the bytes: read 4 bytes with struct.unpack_from(">I", data, offset), then slice exactly that many bytes, and check there really are that many left.