Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation
Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However, the dynamics of adapting these models to low-resource languages with complex morphology and written in unique scripts are not yet explored. Using Tibetan as a case study, we analyze how continued pre-training of pixel-based LMs is influenced by data scale, initial script exposure, and cross-lingual transfer from languages written in other Brahmic scripts. We introduce four ren
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Large Language Models for Small Languages - logos-pres.md →
- PossiblePossibly related (embedding) · 50%Improving machine-translated novels via style transfer — looking for advice on the faithfulness/fluency tradeoff [P] →
- PossiblePossibly related (embedding) · 49%The shrinking landscape of linguistic diversity in the age of large language models - Nature →
- PossiblePossibly related (embedding) · 49%J-Wash: A novel way to brainwash and customize large language models based on Anthropic's Jacobian-Lens! →
- PossiblePossibly related (embedding) · 48%Mortif Technologies' own Large Language Model (LLM) ranked third among open weight models in the glo.. - 매일경제 →
- LinkedLinked via arxiv author · 85%Yiran Zhang →
“Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation”
- LinkedLinked via arxiv author · 85%Miryam de Lhoneux →
“Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation”
- LinkedLinked via arxiv author · 85%Wessel Poelman →
“Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation”
