Learning on Web Dev Open is free for all.

Interview Prep · System DesignDesign a web crawler
← System Design

Design a web crawler

Medium45 minFree, no account

A queue that must not grow forever, and a politeness constraint that shapes the whole design.


The question

Design a crawler that fetches a billion pages a month and hands the content to an indexing pipeline.

It must not get your IP range banned.

Functional
  • Start from a seed set, fetch pages, extract links, continue.
  • Re-crawl pages on a schedule that reflects how often they change.
  • Hand page content to a downstream consumer.
Non-functional
  • 1B pages/month ≈ 400 pages/sec sustained.
  • Never more than one request per second to a single host.
  • Obey robots.txt.
  • A crashed worker must not lose or duplicate work.
45:00Commit to an answer before you open the solution. Reading it first teaches you to recognise good answers, which is not the skill being tested.

Stuck?

0 of 3 hints taken

The worked solution

written by a person · not a grade

Score yourself

0 of 5 marked
  • Partitioned the frontier by host and derived politeness from it25
  • Used a Bloom filter for URL dedup and justified the error direction20
  • Separated URL dedup from content dedup15
  • Made work recoverable with visibility timeouts and at-least-once20
  • Addressed re-crawl frequency and traps20

We run no AI here and nothing on this page grades you. The score is yours, and the useful number is the one you get on the same problem a month from now, cold.

kept in this browser only