newsBAIR (Berkeley)Trust 88 · LabPublished 11mo agoLive · 1mo ago
What exactly does word2vec learn?
What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known precursor to modern language models, for many years, researchers lacked a quantitative and predictive theory describing its learning pro
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownTransformer →
- PossiblePossibly related (embedding) · 47%TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models →
- PossiblePossibly related (embedding) · 47%Complexity-Guided Component-wise Initialization for Language Model Pretraining →
- PossiblePossibly related (embedding) · 50%Scalable Visual Pretraining for Language Intelligence →
- PossiblePossibly related (embedding) · 51%Language Identification with Succinct Machine-Independent Traces →
- PossiblePossibly related (embedding) · 50%Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers →
Covers
Covers (incoming)
paperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperRepresentational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language ModelspaperTypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code ModelspaperComplexity-Guided Component-wise Initialization for Language Model PretrainingpaperScalable Visual Pretraining for Language IntelligencepaperLanguage Identification with Succinct Machine-Independent TracespaperControlling Implicit Shortcut Reliance in L2 Spoken English Auto-markerspaperExposure is Optional: Learning Unlike Coordination in Language ModelspaperWhat, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations
Related across the graph
paperScalable Visual Pretraining for Language IntelligencepaperLanguage Identification with Succinct Machine-Independent TracespaperExposure is Optional: Learning Unlike Coordination in Language Modelsglossary_termTransformerpaperRepresentational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language ModelspaperControlling Implicit Shortcut Reliance in L2 Spoken English Auto-markerspaperTypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code ModelspaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperComplexity-Guided Component-wise Initialization for Language Model PretrainingpaperWhat, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations
