Schema design in a schemaless store
MongoDB does not enforce a schema, which does not mean your data has not got one. It means the schema is in your application code and nobody wrote it down.
Every collection has a schema. The only question is whether it is declared in one place or implied by forty spots in your code that assume a field exists. Schemaless removes the migration step, which is genuinely useful early, and moves the cost to the day you find three generations of document shape in the same collection and every read has to handle all three.
Model around access patterns, not around entities. The relational instinct is to normalise and join later; the document instinct is to store together what you read together. If every page that shows an order also shows its line items and never shows a line item alone, embed them and read one document. If comments on a post number in the thousands and are paginated, they are their own collection with a reference, because embedding them puts you on a collision course with the sixteen megabyte document limit and makes every post read enormous.
The rule that decides most cases: embed when the child is owned by the parent, bounded in size, and always read with it. Reference when the child is shared, unbounded, or has a life of its own. Write the rule you applied in a comment above the model, because in six months the shape will look arbitrary and someone will "fix" it.
You should now be able to
- Design a document shape around access patterns
- Choose between embedding and referencing on stated grounds
- Explain what schemaless actually removes
Loading…