Learning on Web Dev Open is free for all.

AI-Native Products > Retrieval that survives contactEmbeddings without mysticism
Phase 08Retrieval that survives contact404 of 434

Embeddings without mysticism

A vector of numbers where nearby means similar in the way the model was trained to measure, which is not always the way you meant.

Concept14 minAI pair

An embedding maps text to a point in a few hundred or thousand dimensions, positioned so that text with similar meaning lands nearby. Similarity is usually cosine distance, and the resulting number has no absolute meaning: 0.82 is not a percentage of relevance, and the useful threshold differs by model, by domain and by chunk length. Calibrate it on your own data or do not use a threshold at all.

Similar is not the same as relevant, and the gap causes most retrieval disappointment. Negations embed close to their affirmations. Two documents about billing look alike whether one answers the question or contradicts it. Exact identifiers, an error code, a part number, a surname, are exactly what embeddings are worst at, because they carry meaning by identity rather than by semantics.

That last point is the practical reason nobody serious uses vectors alone. Keyword search finds the part number; the vector finds the paraphrase. You need both, which is where hybrid search comes from later in this lesson.

You should now be able to

  • Explain what a similarity score does and does not mean
  • Choose an embedding model on the basis of your own data
Ask the community

Loading…