Complexity-Guided Component-wise Initialization for Language Model Pretraining
Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style language-model pretraining. First, we analyze eleven pretrained GPT-2-style checkpoints that vary in size, language, tokenizer, and training corpus, measuring Frobenius norm and effective-rank entropy across layers and Transformer subcomponents. The checkpoints show shared depth trends, especially increasing scale and strong
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 48%GPT-2 Fully Decoded Internally Black Box Fully Open With Demo →
- PossiblePossibly related (embedding) · 48%Transformer →
- PossiblePossibly related (embedding) · 47%What exactly does word2vec learn? →
- PossiblePossibly related (embedding) · 47%PacificAI/langtest →
- LinkedLinked via arxiv author · 85%Konstantin Garbers →
“Complexity-Guided Component-wise Initialization for Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Nicholas Oh →
“Complexity-Guided Component-wise Initialization for Language Model Pretraining”
