~/problems / Files & networks / File deduplication

Basics: group paths by identical content

easy basics ~10 min

Before walking real directories, get the core grouping right on data that's already in memory.

Implement group_duplicates(files: dict[str, bytes]) -> list[list[str]]. files maps a path to that file's contents.

  • Return every group of two or more paths whose contents are exactly equal. A path whose content is unique doesn't appear.
  • Sort the paths inside each group, then sort the list of groups.
  • Empty contents (b"") are equal to each other like any other value.
  • Identify content by its SHA-256 digest (hashlib.sha256(data).digest()): in the real problem you can't keep every file in memory, but you can keep one small hash per file.
group_duplicates({
    "a.txt": b"hello",
    "x/b.txt": b"hello",
    "c.txt": b"hellp",
    "d": b"",
    "e": b"",
})
# [["a.txt", "x/b.txt"], ["d", "e"]]

group_duplicates({"only": b"data"})   # []
Show hint

build a defaultdict(list) from digest to paths, then keep only the lists with at least two paths.

Topic: File deduplication. Group by size, then hash; walk directories.

0:00
Ctrl ' run · Ctrl ↵ submit
esc