~/problems
Problems
Basics first, then the classics, then company-style assessments. Every problem has tests you run right here; multi-level ones unlock as you go. See the roadmap.
File deduplication
Group by size, then hash; walk directories.
Notes
Recognise it when: "find duplicate files", "detect identical content" at scale.
Pipeline: filter by size (free: os.path.getsize), then maybe hash the first 4 KB, then a full hash, all in chunks.
h = hashlib.sha256()
with open(p, "rb") as f:
while chunk := f.read(65536): h.update(chunk)
Talking points: never read a file whose size is unique. Streaming keeps memory flat. Skip symlinks and special files. Hash collisions: compare bytes if you're paranoid. Parallelize the hashing, which is I/O-bound.
3 problems
Practical systems
Files & networks Deduplicating files, resolving names.
File deduplication
- Basics: group paths by identical content basics easy
- Clean up notes synced from two laptops easy
- Find duplicate files (size, prefix, content) 3 levels Anthropic medium