newsIEEE Spectrum AITrust 88 · LabPublished 2mo agoLive · 1mo ago
New Server Hopes to Break Through AI’s “Memory Wall”
Memory is arguably the most serious constraint on modern AI large language models (LLMs). According to one influential paper , LLM token generation is an inherently memory-bound task, meaning the rate at which models output text is lim
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSparse attention at million-token context →
- LinkedLinked via unknownNoshkoto/Noshy →
- LinkedLinked via unknownCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention →
- LinkedLinked via unknownSpeculative decoding with draft models →
- LinkedLinked via unknownScaling limit of the Random Language Model →
- LinkedLinked via unknownFrom Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond →
- LinkedLinked via unknownSelective Memory Retention for Long-Horizon LLM Agents →
Covers
Covers (incoming)
paperScaling limit of the Random Language ModelpaperFrom Tokens to States: LLMs as a Special Case of World Models and the Continuous Path BeyondpaperSelective Memory Retention for Long-Horizon LLM AgentspaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperEvolution Fine-Tuning: Learning to Discover Across 371 Optimization TaskspaperMulti-Block Diffusion Language ModelspaperRepresentational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language ModelspaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperSurrogate Fidelity: When Can Open LLMs Explain Closed Ones?paperCHERRY: Compressed Hierarchical Experts with Recurrent Representational YieldpaperUnderstanding Large Language Modelsrepoplur-ai/plurrepomem0ai/mem0repoAI-Hypercomputer/maxtextrepomanojmallick/sigmaprepojordanhubbard/nanolangrepogaran0613/ai-memory-gatewayrepollmsresearch/llm-flashcardsrepojaylfc/taosmdpaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperWhen Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?paperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperA Practical Investigation of Training-free Relaxed Speculative Decodingrepogdevenyi/huggingface-estimaterepoNovasPlace/CSMrepoScottcjn/ram-cofferspaperSigLIP-HD by Fine-to-Coarse SupervisionpaperNeural Collapse Is Forbidden: Information Floors in Language ModelspaperProduction and Perception in LLMs: A Token Probability ApproachpaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionpaperIn-Place Tokenizer Expansion for Pre-trained LLMspaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic ContentpaperAdaptive Multi-Step Lookahead Decoding for Diffusion Language ModelspaperSelectInfer: Selective Neuron Loading and Computation for On-Device LLMspaperHow Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-TuningpaperWindowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Contextrepodiegosouzapw/OmniGlyph
Related across the graph
paperProduction and Perception in LLMs: A Token Probability ApproachpaperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperSelectInfer: Selective Neuron Loading and Computation for On-Device LLMspaperEvolution Fine-Tuning: Learning to Discover Across 371 Optimization TaskspaperMulti-Block Diffusion Language ModelspaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scalingrepomem0ai/mem0paperCHERRY: Compressed Hierarchical Experts with Recurrent Representational Yieldrepogaran0613/ai-memory-gatewaypaperWhen Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?paperRepresentational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language ModelsrepoScottcjn/ram-cofferspaperNeural Collapse Is Forbidden: Information Floors in Language ModelspaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Contentrepodiegosouzapw/OmniGlyphpaperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear AttentionpaperIn-Place Tokenizer Expansion for Pre-trained LLMspaperHow Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-TuningpaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperSigLIP-HD by Fine-to-Coarse Supervisionrepollmsresearch/llm-flashcardspaperSparse attention at million-token contextpaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inferencerepogdevenyi/huggingface-estimatepaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperSurrogate Fidelity: When Can Open LLMs Explain Closed Ones?repoNovasPlace/CSMpaperUnderstanding Large Language ModelspaperAdaptive Multi-Step Lookahead Decoding for Diffusion Language ModelspaperSelective Memory Retention for Long-Horizon LLM AgentspaperScaling limit of the Random Language ModelpaperA Practical Investigation of Training-free Relaxed Speculative Decodingrepojordanhubbard/nanolangrepojaylfc/taosmdrepoplur-ai/plurrepoAI-Hypercomputer/maxtextpaperFrom Tokens to States: LLMs as a Special Case of World Models and the Continuous Path BeyondrepoNoshkoto/Noshyrepomanojmallick/sigmappaperSpeculative decoding with draft modelspaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionpaperWindowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
