~/problems

Problems

Basics first, then the classics, then company-style assessments. Every problem has tests you run right here; multi-level ones unlock as you go. See the roadmap.

Files & networks

File deduplication

Group by size, then hash; walk directories.

Notes

Recognise it when: "find duplicate files", "detect identical content" at scale.

Pipeline: filter by size (free: os.path.getsize), then maybe hash the first 4 KB, then a full hash, all in chunks.

h = hashlib.sha256()
with open(p, "rb") as f:
    while chunk := f.read(65536): h.update(chunk)

Talking points: never read a file whose size is unique. Streaming keeps memory flat. Skip symlinks and special files. Hash collisions: compare bytes if you're paranoid. Parallelize the hashing, which is I/O-bound.

3 problems

Practical systems

Files & networks Deduplicating files, resolving names.

esc