Word embedding
A word embedding represents a word as a dense vector of numbers. Static methods such as word2vec learn one vector per vocabulary word from patterns of use, so related words can have similar representations.
A simple review classifier could look up vectors for each word, average them and use the resulting vector to predict positive or negative sentiment. It can learn associations between words such as “excellent” and “great” without treating every spelling as an unrelated input. Averaging, however, loses word order and can blur a negation such as “not great.”
In a static table, “bank” gets the same vector in “river bank” and “bank account.” A contextual model instead computes a representation for each occurrence using surrounding text. Word vectors therefore differ from the sentence or passage embeddings used by many retrieval systems.
This matters when selecting a representation. A compact word table may suit a simple classifier, while a task that depends on word sense or sentence meaning needs more context. Vector similarity is learned from the corpus; it does not establish factual equivalence.
Sources
- Mikolov et al.: Efficient Estimation of Word Representations — Introduces continuous word representations and evaluates syntactic and semantic similarity.
Go deeper
- Stanford: GloVe docs
Explore a different word-vector objective, pretrained tables and code based on global co-occurrence statistics.