newsGoogle News — LLMTrust 62 · AggregatorPublished 1mo agoLive · 1mo ago
Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%PacificAI/langtest →
- PossiblePossibly related (embedding) · 58%How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation →
- PossiblePossibly related (embedding) · 55%EuroEval/EuroEval →
- PossiblePossibly related (embedding) · 54%Jeryi-Sun/LLM-and-Law →
- PossiblePossibly related (embedding) · 54%Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities →
- PossiblePossibly related (embedding) · 54%Furyton/awesome-language-model-analysis →
- PossiblePossibly related (embedding) · 56%AI4Finance-Foundation/FinGPT →
- PossiblePossibly related (embedding) · 60%piesauce/awesome-dLLM-resources →
Covers
repoPacificAI/langtestpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationrepoEuroEval/EuroEvalrepoJeryi-Sun/LLM-and-LawpaperData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data ComplexitiesrepoFuryton/awesome-language-model-analysis
Covers (incoming)
repoAI4Finance-Foundation/FinGPTrepopiesauce/awesome-dLLM-resourcespaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Contentrepofacebookresearch/stopespaperAgainst Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Mediarepozhangxjohn/LLM-Agent-Benchmark-ListrepoMaartenGr/BERTopicpaperFlesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic ModelspaperCross-lingual Biography Enrichment via Claim Extraction and AlignmentpaperKyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understandingpapernews-crawler-LM: A Small Long-Context Model For High-Quality News CrawlingpaperWhen Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMsrepomohsenhariri/scoriopaperBabelSteering: Multilingual Safety Alignment via English Steering VectorspaperHealMed: Multilingual Evaluation of Large Language Models in MedicinepaperAuditing Cross-Lingual Fairness in Language Model WatermarkingpaperThe Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce LanguagepaperLost in Speech: Trilingual Spoken Hallucination Detection Across Audio and TranscriptspaperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagespaperThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelspaperGraded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal SteeringpaperPretraining Data Can Be Poisoned through Computational PropagandapaperAutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated JournalismpaperDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Related across the graph
paperCross-lingual Biography Enrichment via Claim Extraction and AlignmentpaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperHealMed: Multilingual Evaluation of Large Language Models in Medicinerepofacebookresearch/stopesrepoJeryi-Sun/LLM-and-LawpaperAuditing Cross-Lingual Fairness in Language Model WatermarkingpaperAgainst Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social MediapaperGraded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal SteeringpaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic ContentrepoMaartenGr/BERTopicpaperFlesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic ModelspaperPretraining Data Can Be Poisoned through Computational Propagandapapernews-crawler-LM: A Small Long-Context Model For High-Quality News CrawlingpaperDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training DatapaperBabelSteering: Multilingual Safety Alignment via English Steering VectorspaperThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelspaperAutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated JournalismpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperWhen Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMspaperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagesrepoPacificAI/langtestpaperKyrgyzLLM-Bench: Benchmarking Kyrgyz Language UnderstandingrepoAI4Finance-Foundation/FinGPTpaperThe Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Languagerepopiesauce/awesome-dLLM-resourcesrepozhangxjohn/LLM-Agent-Benchmark-ListrepoEuroEval/EuroEvalrepomohsenhariri/scoriopaperData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data ComplexitiespaperLost in Speech: Trilingual Spoken Hallucination Detection Across Audio and TranscriptsrepoFuryton/awesome-language-model-analysis
