Learning on Web Dev Open is free for all.

Interview PrepML System Design
← Practice

ML System Design

7h of practice10 problems

Machine-learning design rounds are lost at the framing stage far more often than at the modelling stage. Candidates reach for an architecture before anyone has said what the label is, what the baseline would be, or how you would know the thing is working a month after launch.

Every problem here is worked in the same order: the product goal, the prediction, the data and the label, the baseline, the features, the model, serving, and the feedback loop that will eventually poison it.

The round · Forty-five to sixty minutes. Expect the first ten on framing, what exactly is being predicted, from what, and what metric decides, and the rest on data, serving, and what goes wrong in production.

Difficulty
Topic

10 problems

Medium

5
  1. Design feed rankingThe label is the hard part, and optimising the obvious one makes the product worse.Ranking · Multi-objective · Feedback loops45 min
  2. Design search rankingRetrieval and ranking are different problems, and relevance has no ground truth.Retrieval · Learning to rank · Evaluation45 min
  3. Design a recommendation homepageNo query, no intent, and a cold-start problem on both sides.Candidate generation · Embeddings · Cold start45 min
  4. Design ETA predictionA regression problem where being early and being late cost different amounts.Regression · Asymmetric loss · Real-time features40 min
  5. Design demand forecastingA forecast whose errors cost different amounts in each direction, evaluated by a split nobody gets right.Time series · Asymmetric cost · Hierarchies · Backtesting45 min

Hard

5
  1. Design a fraud detection systemOne in a thousand positives, an adversary who adapts, and a false positive that costs a customer.Imbalanced data · Thresholds · Adversarial drift45 min
  2. Design content moderationThe errors are asymmetric, the policy changes weekly, and humans are part of the architecture.Multi-modal · Human in the loop · Policy45 min
  3. Design click-through rate predictionThe one problem where a well-ranked but badly calibrated model costs real money.Calibration · High cardinality · Online learning45 min
  4. Design a retrieval-augmented assistantEasy to demo, hard to evaluate, and the failure mode is a confident wrong answer with a citation.Retrieval · Evaluation · Grounding45 min
  5. Design ML model monitoringThe model is good on launch day. This is the system that tells you when it stops being.Drift · Observability · Delayed labels · Rollout45 min