When Can We Work in Embedding Space? What Text Embeddings Preserve
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers →
- PossiblePossibly related (embedding) · 46%Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers →
- PossiblePossibly related (embedding) · 46%Embedding →
- FuzzySimilar title/name (fuzzy) · 59%FlagOpen/FlagEmbedding →
“Fuzzy title match (0.73): “When Can We Work in Embedding Space? What Text Embeddings Pr” ≈ “FlagOpen/FlagEmbedding””
- LinkedLinked via arxiv author · 85%Simon Freyaldenhoven →
“When Can We Work in Embedding Space? What Text Embeddings Preserve”
