glossary termAngestromTrust 60Published 2mo agoLive · 2mo ago
Transformer
The neural network architecture behind most modern language models.
The neural network architecture behind most modern language models. The neural network architecture behind most modern language models.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownengineering87/llm-atlas →
- LinkedLinked via unknownsentence-transformers/all-MiniLM-L6-v2 →
- LinkedLinked via unknownWhen are likely answers right? On Sequence Probability and Correctness in LLMs →
- LinkedLinked via unknownWhat exactly does word2vec learn? →
- LinkedLinked via unknownIdentifying Interactions at Scale for LLMs →
- LinkedLinked via unknownminimal-diffusion-lm →
- LinkedLinked via unknownvlm-starter →
Related to (incoming)
repoengineering87/llm-atlasmodelsentence-transformers/all-MiniLM-L6-v2paperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigationrepominimal-diffusion-lmrepovlm-starterpaperVASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned AnchoringpaperGeneralization Analysis of Transformers in Distribution RegressionpaperAdaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language ModelspaperUnderstanding Evaluation Illusion in Diffusion Large Language ModelspaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning TaskspaperMorphing into Hybrid Attention ModelspaperCoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty EstimationpaperLOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype AlignmentpaperTeam MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline DynamicspaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionpaperBridging the Gap Between Latent and Explicit Reasoning with Looped TransformerspaperCHERRY: Compressed Hierarchical Experts with Recurrent Representational YieldpaperExplicit Fuzzy Logic in the Feed-Forward Layer: Self-Forgetting Quantifiers Discover Legible Grammatical-Licensing DetectorspaperUnderstanding Large Language ModelspaperThe State-Prediction Separation HypothesispaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Trainingreposgl-project/sglangreporobertknight/rtenpaperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperHaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI SafetypaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperOn the Role of Directionality in Structural GeneralizationpaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperTransformer Geometry Observatory TGO-II: Representational Similarity ObservatoryrepoEuroEval/EuroEvalrepox-tabdeveloping/turftopicrepoIBM/LNNrepoNixtla/neuralforecastrepokyegomez/BitNetrepobitsandbytes-foundation/bitsandbytespaperUmm... With Transformers? Insights from Filled Pause Use across Four Slavic ParliamentspaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperPrompt Compression via Activation AggregationpaperIt Takes a MAESTRO To Prune Bad Expertsrepokesimeg/awesome-turkish-language-modelsrepoaibridge-afrilabs-group/aibridge-afrilabs-projectrepoFuryton/awesome-language-model-analysisrepogrisuno/TopoGPT2paperComplexity-Guided Component-wise Initialization for Language Model PretrainingpaperNeural Collapse Is Forbidden: Information Floors in Language ModelspaperFreyaTTS Technical ReportpaperTokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASRpaperFrom Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASPpaperEncoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language ModelspaperInvariant Learning Dynamics of Transformers in Inductive Reasoning TaskspaperAudio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Modelrepoerogol/BlaGPTrepodlidstrom/NeuralNetworkInAllLangspaperLinear representations of grammaticality in neural language modelspaperT^2MLR: Transformer with Temporal Middle-Layer RecurrencepaperLanguage Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio GeometrypaperInduction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language ModelspaperFrontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2DpaperMobius Learning: Cyclic Depth Folding in TransformerspaperL1 Augmented Attention as an Improved Vector Similarity MetricpaperSelective State-Space Adaptation and Retrieval for Language Model ReasoningpaperExposure is Optional: Learning Unlike Coordination in Language ModelspaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformerspaperWhat, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations
Covers (incoming)
newsWhat exactly does word2vec learn?newsIdentifying Interactions at Scale for LLMsnewsBook Review: Domain-Specific Small Language Models by Guglielmo IozzianewsKnowledge Distillation of Black-Box Large Language ModelsnewsKnowledge Distillation of Black-Box Large Language Models (2024)newsTraining transformers where every layer W = V·Uᵀ from initialization reveals a corpus-determined optimal rank - looking for arXiv endorser (cs.LG) [D]newsTransformers in Deep Learning: How Self-Attention Changed Modern AI - SnowflakenewsDiffusion Models: The Deep Learning Architecture Behind Modern Generative AI - Snowflake
Related across the graph
paperOn the Role of Directionality in Structural GeneralizationpaperFrontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2DpaperTeam MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline DynamicsnewsKnowledge Distillation of Black-Box Large Language ModelsnewsDiffusion Models: The Deep Learning Architecture Behind Modern Generative AI - SnowflakepaperEncoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language ModelspaperInduction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language ModelspaperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperThe State-Prediction Separation HypothesispaperExposure is Optional: Learning Unlike Coordination in Language ModelspaperFrom Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASPpaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperUnderstanding Evaluation Illusion in Diffusion Large Language ModelspaperCHERRY: Compressed Hierarchical Experts with Recurrent Representational YieldpaperL1 Augmented Attention as an Improved Vector Similarity MetricpaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformerspaperFreyaTTS Technical ReportnewsTraining transformers where every layer W = V·Uᵀ from initialization reveals a corpus-determined optimal rank - looking for arXiv endorser (cs.LG) [D]repominimal-diffusion-lmrepobitsandbytes-foundation/bitsandbytespaperT^2MLR: Transformer with Temporal Middle-Layer RecurrencepaperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasksrepoengineering87/llm-atlaspaperExplicit Fuzzy Logic in the Feed-Forward Layer: Self-Forgetting Quantifiers Discover Legible Grammatical-Licensing DetectorspaperIt Takes a MAESTRO To Prune Bad ExpertspaperCoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimationrepokyegomez/BitNetpaperNeural Collapse Is Forbidden: Information Floors in Language ModelspaperLanguage Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio GeometrypaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperMorphing into Hybrid Attention Modelsmodelsentence-transformers/all-MiniLM-L6-v2paperTransformer Geometry Observatory TGO-II: Representational Similarity Observatoryrepox-tabdeveloping/turftopicpaperInvariant Learning Dynamics of Transformers in Inductive Reasoning TaskspaperMobius Learning: Cyclic Depth Folding in TransformerspaperPrompt Compression via Activation AggregationpaperVASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned AnchoringnewsWhat exactly does word2vec learn?paperUmm... With Transformers? Insights from Filled Pause Use across Four Slavic ParliamentspaperAudio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Modelreposgl-project/sglangpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationnewsBook Review: Domain-Specific Small Language Models by Guglielmo IozziapaperUnderstanding Large Language Modelsrepogrisuno/TopoGPT2paperComplexity-Guided Component-wise Initialization for Language Model Pretrainingrepokesimeg/awesome-turkish-language-modelsnewsKnowledge Distillation of Black-Box Large Language Models (2024)repoaibridge-afrilabs-group/aibridge-afrilabs-projectpaperAdaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language Modelsrepoerogol/BlaGPTpaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperWhat, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representationsreporobertknight/rtenpaperBridging the Gap Between Latent and Explicit Reasoning with Looped Transformersrepodlidstrom/NeuralNetworkInAllLangspaperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperLOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype AlignmentpaperLinear representations of grammaticality in neural language modelsrepoIBM/LNNnewsTransformers in Deep Learning: How Self-Attention Changed Modern AI - SnowflakepaperSelective State-Space Adaptation and Retrieval for Language Model ReasoningpaperHaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI SafetyrepoNixtla/neuralforecastrepoEuroEval/EuroEvalpaperGeneralization Analysis of Transformers in Distribution RegressionpaperTone-Conditioned Curriculum Learning for Low-Resource Bantu Speech RecognitionrepoFuryton/awesome-language-model-analysispaperTokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASRrepovlm-starternewsIdentifying Interactions at Scale for LLMs
