Someone synced their plain-text notes between a Windows laptop and a Mac for years. Now the notes folder holds many copies of the same note that differ only in line endings or in how many blank lines trail at the end. Plan the cleanup: which copy to keep, and which to delete.
Implement cleanup_plan(files: dict[str, bytes]) -> dict[str, str]. files maps a path to its raw contents.
Two notes count as the same when their normalized contents are equal. To normalize:
- replace every
b"\r\n"withb"\n"(a loneb"\r"stays as it is), then - remove all
b"\n"bytes at the very end.
So b"milk\r\neggs\r\n", b"milk\neggs" and b"milk\neggs\n\n\n" are the same note, and b"", b"\n" and b"\r\n" are all the same (empty) note. Trailing spaces are not stripped.
In each group of two or more same notes, keep the path with the fewest characters (ties: the alphabetically smallest). The result maps every other path in the group to the path that's kept. Paths that have no duplicate, and kept paths, don't appear in the result.
cleanup_plan({
"notes/shopping.txt": b"milk\r\neggs\r\n",
"mac/notes/shopping.txt": b"milk\neggs",
"s.txt": b"milk\neggs\n\n",
"todo.txt": b"call mum\n",
"todo copy.txt": b"call mum \n", # trailing space: a different note
})
# {"notes/shopping.txt": "s.txt", "mac/notes/shopping.txt": "s.txt"}
As in Basics, group by a SHA-256 digest (of the normalized bytes) rather than comparing every pair of files: there can be tens of thousands of notes.
Show hint
data.replace(b"\r\n", b"\n").rstrip(b"\n") normalizes. Group paths in a defaultdict(list) keyed by the digest; for each group of size ≥ 2 pick the keeper with min(paths, key=lambda p: (len(p), p)).