repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 23d ago
thu-pacman/chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference →
- PossiblePossibly related (embedding) · 59%Understanding Large Language Models →
- PossiblePossibly related (embedding) · 58%When are likely answers right? On Sequence Probability and Correctness in LLMs →
- PossiblePossibly related (embedding) · 58%Message Passing Enables Efficient Reasoning →
- PossiblePossibly related (embedding) · 55%Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian →
- PossiblePossibly related (embedding) · 46%The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal Dependencies →
- PossiblePossibly related (embedding) · 49%NAVER LABS Europe Submission to the Instruction-following 2026 Short Track →
- PossiblePossibly related (embedding) · 51%Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation →
Implements
paperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperUnderstanding Large Language ModelspaperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperMessage Passing Enables Efficient ReasoningpaperCross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian
Implements (incoming)
paperThe Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal DependenciespaperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperBayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty EstimationpaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperBamiBERT: A New BERT-based Language Model for VietnamesepaperOn the Role of Directionality in Structural GeneralizationpaperAn Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and GenerationpaperReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningpaperHeaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language ModelspaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperEstimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMspaperData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data ComplexitiespaperResample or Reroute? Budget-Aware Test-Time Model Selection for Large Language ModelspaperEchoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-NungpaperIt Takes a MAESTRO To Prune Bad ExpertspaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingpaperThe Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMspaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperSelf-Guided Test-Time Training for Long-Context LLMspaperA Sovereign, Open-Source Foundation Model for German and EnglishpaperTokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASRpaperExtending LLM Context via Associative Recurrent MemorypaperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagespaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-ExpertspaperAccelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference TechniquespaperThe Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
Covers (incoming)
Related across the graph
paperOn the Role of Directionality in Structural GeneralizationpaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelspaperAccelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference TechniquespaperReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningpaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperAn Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and GenerationpaperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperEstimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMspaperCross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on RomanianpaperEchoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-NungpaperExtending LLM Context via Associative Recurrent MemorypaperResample or Reroute? Budget-Aware Test-Time Model Selection for Large Language ModelspaperIt Takes a MAESTRO To Prune Bad ExpertspaperThe Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMspaperHeaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language ModelspaperMessage Passing Enables Efficient ReasoningpaperBayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty EstimationpaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingpaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperThe Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant ContextpaperA Sovereign, Open-Source Foundation Model for German and EnglishnewsIf you had a 300M parameter model, what would you optimize it for?paperThe Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal DependenciespaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperUnderstanding Large Language ModelspaperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagespaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperBamiBERT: A New BERT-based Language Model for VietnamesepaperNAVER LABS Europe Submission to the Instruction-following 2026 Short TrackpaperSelf-Guided Test-Time Training for Long-Context LLMspaperData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data ComplexitiespaperTokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASRpaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
