PlainLogic

Interactive lab · Practical AI

Embeddings: Meaning as Numbers

An embedding turns a word, sentence, or image into a list of numbers — and suddenly similarity becomes distance on a map.

Plain-language AI

In plain logic

No real AI runs on this page. Everything is explained and illustrated in your browser — no model calls, no accounts, nothing sent anywhere.

An embedding model maps text, images, or other inputs to vectors — long lists of numbers. Items used in similar contexts tend to land near each other on the map. Real embeddings have hundreds or thousands of dimensions; humans can only look at two or three, so visualizations squash the real space down and show the rough neighborhoods.

Search systems compare these vectors to find candidates that may be relevant even when the words differ. Ask for “cheap evening meal” and an embedding search can surface a page about “affordable dinner” that shares zero keywords with your query. That meaning-level matching is what powers modern semantic search — and the retrieval half of RAG.

more about food more about travel pizza sandwich soup car train airplane bicycle query: "quick lunch near the station" nearest: sandwich bicycle train

Simplified educational illustration — hand-placed 2D points, not real model output. Real embeddings have hundreds of dimensions and are learned, not assigned by hand.

The dashed lines tell the story: the query “quick lunch near the station” sits between the food cluster and the travel cluster, and its nearest neighbors mix both worlds — sandwich, bicycle, train. Nearness pulled in meaning from both sides of the question, even though the query shares no exact words with any point.

Hands-on

Try this

Our map uses hand-assigned food and transport scores — a toy version of the real thing. Try this thought experiment: where would you place “airport sandwich”? High on both axes, right in the middle. Where would “tire” go? Far toward travel, zero toward food.

Now imagine doing that for every word in a library, in a thousand dimensions, with positions learned from billions of sentences instead of assigned by hand. That is a real embedding model — and the reason “cheap evening meal” can find “affordable dinner” without sharing a single word.

Honest boundaries

What this leaves out

Nearness is a clue, not proof of equal meaning or truth. Two sentences can sit side by side because they share a topic while disagreeing completely (“vaccines work” vs “vaccines don't work” can embed closely — they are about the same thing). Embeddings capture aboutness, not correctness.

The model, the distance measure, and the task all change which results are useful. Embeddings also inherit the biases of their training data — whatever associations the text contained, the map quietly preserves. Use them for candidate retrieval, then verify with the actual text.

Honest answers

Questions people ask

What is a vector, really?

Just a list of numbers — like coordinates. “Pizza” might be [0.21, -0.87, 0.44, …] with hundreds of entries. The numbers have no individual meaning; only the relative positions of points matter.

Why can't we just see the real map?

Because it has hundreds of dimensions and human eyes do three. Tools like t-SNE squash it down for visualization, but the squashed picture distorts distances — treat 2D embedding plots as rough sketches, never measurements.

Do embeddings understand truth?

No. They capture how words are used, not whether statements are true. “The moon is cheese” embeds near moon-talk, not near falsehood — there is no truth dimension on the map.