An embedding is a way of turning a piece of text, an image, or another piece of content into a list of numbers that captures its meaning, so that two pieces of content with similar meaning end up with similar numbers — even if they do not share a single word in common.
OpenAI’s own developer documentation defines the term plainly: “an embedding is a vector (list) of floating point numbers. The distance between two vectors measures their relatedness. Small distances suggest high relatedness and large distances suggest low relatedness”. That distance, not any shared keyword, is what a system checks when it decides whether two pieces of content mean roughly the same thing.
That is the same mechanism behind a searchable knowledge base: the search compares embeddings, not the words themselves, which is why a knowledge base can surface the right answer even when the question is phrased nothing like the source document.
A client who asks whether they can get out of their mortgage early, and another who asks what happens if they break their term, share almost no words in common, but they are asking the same underlying question. Their embeddings — the numeric form each question is turned into — sit close together for exactly that reason, which is what lets a support tool treat them as the same question even though a simple keyword search would never connect the two. Nothing about that closeness is inspected or explained by a person; it falls directly out of the list of numbers each question was converted into, which is why an embedding-based search can occasionally group two questions that a human reader would not, with no obvious record of why.
See also: what is a context window, what is natural language processing, what is machine learning.
Choosing what gets turned into embeddings, and how often that index is refreshed, is a build decision — custom-ai-solutions covers what keeps a search tool like this accurate as the underlying material changes.