~/problems / Pools & pipelines / Thread pool / concurrent crawler

Basics: parallel map with a pool

easy basics ~10 min

Implement parallel_map(fn, items, max_workers) -> list:

  • Call fn(item) for every item, using a pool of at most max_workers threads so the slow calls (think: HTTP requests) overlap.
  • Return the results in the same order as items, even though the calls finish in any order.
  • If any call raises an exception, parallel_map raises it too.
parallel_map(lambda x: x * x, [3, 1, 2], 2)   # [9, 1, 4]
parallel_map(str, [], 4)                      # []

If fn sleeps 0.1s, then 8 items with max_workers = 4 take about 0.2s instead of 0.8s.

Constraints: max_workers >= 1, 0 <= len(items) <= 100.

Show hint

with ThreadPoolExecutor(max_workers=...) as pool: then pool.map(fn, items) yields results in input order and re-raises a worker's exception when you reach its result (pool.submit plus future.result() works too).

Topic: Thread pool / concurrent crawler. ThreadPoolExecutor, asyncio, thread-safe visited set.

0:00
Ctrl ' run · Ctrl ↵ submit
esc