Design a web crawler
Medium45 minFree, no account
A queue that must not grow forever, and a politeness constraint that shapes the whole design.
The question
Design a crawler that fetches a billion pages a month and hands the content to an indexing pipeline.
It must not get your IP range banned.
Functional
- Start from a seed set, fetch pages, extract links, continue.
- Re-crawl pages on a schedule that reflects how often they change.
- Hand page content to a downstream consumer.
Non-functional
- 1B pages/month ≈ 400 pages/sec sustained.
- Never more than one request per second to a single host.
- Obey robots.txt.
- A crashed worker must not lose or duplicate work.
45:00Commit to an answer before you open the solution. Reading it first teaches you to recognise good answers, which is not the skill being tested.
Stuck?
0 of 3 hints takenThe worked solution
written by a person · not a gradeScore yourself
0 of 5 marked- Partitioned the frontier by host and derived politeness from it25
- Used a Bloom filter for URL dedup and justified the error direction20
- Separated URL dedup from content dedup15
- Made work recoverable with visibility timeouts and at-least-once20
- Addressed re-crawl frequency and traps20
We run no AI here and nothing on this page grades you. The score is yours, and the useful number is the one you get on the same problem a month from now, cold.
kept in this browser only