Read original ↗
newsGoogle News — LLMTrust 62 · AggregatorPublished 1mo agoLive · 1mo ago

Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

paperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagespaperThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelspaperKyrgyzLLM-Bench: Benchmarking Kyrgyz Language UnderstandingpaperAutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalismpapernews-crawler-LM: A Small Long-Context Model For High-Quality News CrawlingpaperWhen Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMspaperDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training DatapaperGraded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal SteeringpaperPretraining Data Can Be Poisoned through Computational PropagandarepoAI4Finance-Foundation/FinGPTrepopiesauce/awesome-dLLM-resourcespaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic ContentpaperFlesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic ModelspaperCross-lingual Biography Enrichment via Claim Extraction and AlignmentpaperThe Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce LanguagepaperLost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcriptsrepofacebookresearch/stopesrepomohsenhariri/scoriopaperBabelSteering: Multilingual Safety Alignment via English Steering VectorspaperAgainst Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Mediarepozhangxjohn/LLM-Agent-Benchmark-ListpaperHealMed: Multilingual Evaluation of Large Language Models in MedicinepaperAuditing Cross-Lingual Fairness in Language Model WatermarkingrepoMaartenGr/BERTopic

Related across the graph

paperCross-lingual Biography Enrichment via Claim Extraction and AlignmentpaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperHealMed: Multilingual Evaluation of Large Language Models in Medicinerepofacebookresearch/stopesrepoJeryi-Sun/LLM-and-LawpaperAuditing Cross-Lingual Fairness in Language Model WatermarkingpaperAgainst Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social MediapaperGraded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal SteeringpaperRate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic ContentrepoMaartenGr/BERTopicpaperFlesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic ModelspaperPretraining Data Can Be Poisoned through Computational Propagandapapernews-crawler-LM: A Small Long-Context Model For High-Quality News CrawlingpaperDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training DatapaperBabelSteering: Multilingual Safety Alignment via English Steering VectorspaperThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelspaperAutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated JournalismpaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperWhen Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMspaperTranslation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource LanguagesrepoPacificAI/langtestpaperKyrgyzLLM-Bench: Benchmarking Kyrgyz Language UnderstandingrepoAI4Finance-Foundation/FinGPTpaperThe Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Languagerepopiesauce/awesome-dLLM-resourcesrepozhangxjohn/LLM-Agent-Benchmark-ListrepoEuroEval/EuroEvalrepomohsenhariri/scoriopaperData Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data ComplexitiespaperLost in Speech: Trilingual Spoken Hallucination Detection Across Audio and TranscriptsrepoFuryton/awesome-language-model-analysis