Before walking real directories, get the core grouping right on data that's already in memory.
Implement group_duplicates(files: dict[str, bytes]) -> list[list[str]]. files maps a path to that file's contents.
- Return every group of two or more paths whose contents are exactly equal. A path whose content is unique doesn't appear.
- Sort the paths inside each group, then sort the list of groups.
- Empty contents (
b"") are equal to each other like any other value. - Identify content by its SHA-256 digest (
hashlib.sha256(data).digest()): in the real problem you can't keep every file in memory, but you can keep one small hash per file.
group_duplicates({
"a.txt": b"hello",
"x/b.txt": b"hello",
"c.txt": b"hellp",
"d": b"",
"e": b"",
})
# [["a.txt", "x/b.txt"], ["d", "e"]]
group_duplicates({"only": b"data"}) # []
Show hint
build a defaultdict(list) from digest to paths, then keep only the lists with at least two paths.