Ge Lan
Ge Lan — researcher or builder tracked in the Angestrom contributor network.
Papers · 19
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restric
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focused evaluations rely on synthetic questions and rarely report item-level properties. This paper introduces EduArt, an educational-level benchmark for art-historical knowledge and visual reasoning in multimodal LLMs. EduArt comprises 871 human-authored questions from Italian secondary-school exercises and US Advanced Placement Art History exams, spanning two languages and seven formats from multiple choice to in-text w
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal populati
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards: substantial similarity, which extends to stylistic choices, narrative structure, and creative elaboration. This mismatch between what current methods detect and what the law protects leaves a significant compliance gap. We introduce PSALM, an LLM-as-a-judge framework that operationalises EU copyright doctrine through ten evaluators assessing computational overlap, styli
Understanding Large Language Models
Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing questions about their mechanisms, capabilities, and relationship to human cognition remain highly debated. This chapter aims to outline our current understanding of LLMs by discussing recent evidence on emerging capabilities and their mechanistic implementation within processing layers. We begin with a concise overview of the Transformer architecture, emphasizing how the attention mechanism enables training on massive datasets, allowing LLMs to
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical environments.
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, most existing methods rely on conventional deep learning architectures, and the limited model capacity constrains performance. Although large-model post-training techniques have achieved great success in general domains, their direct transfer to RSICC remains challenging due to data scarcity and the need for fine-grained change understanding. To address this, we propose RSICCLLM, the first post-training framework for la
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selectively adopt relevant information while rejecting deceptive or harmful content is therefore critical for reliable deployment in real-world retrieval settings. We introduce SelectBench, a controlled benchmark and training set for selective evidence adoption, and post-train Qwen3.5-4B directly with DAPO using either dete
News · 22
Mortif Technologies' own Large Language Model (LLM) ranked third among open weight models in the glo.. - 매일경제
<a href="https://news.google.com/rss/articles/CBMiS0FVX3lxTFBWcGZheDRDaHFyV2JlSVZPR0std3VZNDRROWRaMU43VW5CSU1IYlFMQjR3eGZLdnl6MjZPTFB4LWVKNkd5UW1JZWowUQ?oc=5" target="_blank">Mortif Technologies' own Large Language Model (LLM) ranked third among open weight models in the glo..</a> <font color="#6f6f6f">매일경제</font>
US accuses Chinese AI firm Moonshot of stealing tech and using banned Nvidia chips for its large language model. - Pluang
<a href="https://news.google.com/rss/articles/CBMikAFBVV95cUxNbzF0LS1HaEVEaTlCZE9qSGZPc2RuRzY1QlFFNnFzNUZha3N6VnFLREp4d2ZKdnJYT3ZCQ2hiT1hpYnZEY1d0VEZOUFNuUGtQckV1ZTY1ckVNMmE3V0xlaFJERnFEQkY2ZDRCNVhqWGR4Z0IzT1hkNkZFVjZTRmVreHd4Wm1YNENXOFRNMG9tbEw?oc=5" target="_blank">US accuses Chinese AI firm Moonshot of stealing tech and using banned Nvidia chips for its large language model.</a> <font color="#6f6f6f">Pluang</font>
Towards principled knowledge editing methods for large language model reasoning
<p>Nature Machine Intelligence, Published online: 14 August 2026; <a href="https://www.nature.com/articles/s42256-026-01276-y">doi:10.1038/s42256-026-01276-y</a></p>Chen et al. explore limitations of current knowledge editing techniques in large language models and propose three promising research directions that respect the complexity of knowledge representation in a real-world setting.
Large Language Models (LLMs): Transforming Enterprise AI Development - nerdbot
<a href="https://news.google.com/rss/articles/CBMingFBVV95cUxPOEk5MzZoOU1MeFZzWXVDYzJxbm5xR1JTbXV2V0owSWtLWmVZOHh5QTRoQ3FPYUdQVzBoWTRtZHc5blV6RDhKTHFRVmx6T2hSOUVaME5PdGdTaDM4WHZ5cDJJSWJ4b3hva3AzOC1QQXR6bjhRcmtvT2t3RC1nbEdVSnU1Tzl4XzFqa3JpdUNJdExiVnFxdXBwaDBsQjRkUQ?oc=5" target="_blank">Large Language Models (LLMs): Transforming Enterprise AI Development</a> <font color="#6f6f6f">nerdbot</font>
Large Language Models Misinformation: Ecosystem Security Shift - The Cryptonomist
<a href="https://news.google.com/rss/articles/CBMigAFBVV95cUxPd25BTTRhRERoYUVUakh0U1JIaUo5cVZHOGFpaXdnT1hzQmpRZ2ZmTmM3czhoX3NOWGQ4RGYtMEw5UzBBdThlby1wakpQaTZkM05UTjFjMWM5VjBrYnFfemI5WjVINlF6dURsci1BblVqblFCODNpa01uVm1FdWUycQ?oc=5" target="_blank">Large Language Models Misinformation: Ecosystem Security Shift</a> <font color="#6f6f6f">The Cryptonomist</font>
13 Assessing the Quality and Reliability of Online Health Information for Breast Pathology: Search Engines Versus Large Language Models - CancerNetwork
<a href="https://news.google.com/rss/articles/CBMi0wFBVV95cUxNLVZ3TjZNNGg2dXZreFQteS1feFMzQmE0VUJXSGtKMk9ISzhPNUhXcFN3TENOb2hXVGlFRzlmalc3LXRWdDRLMl9KWGc2eTY1OWwzQTBVMkxSeTZKLXplVDdCSHNQZlEzVDhjUlZYdW1RX29jNTNRU0R5M0U4R2J3MnIzb0VMeWk0cDB5UjM1YUdIWUI3YU5hY3REdloyS2NRWEVCWF9XQWp5T2lOck1UU1ZhcmNnQW9WMzB5Z1AxQy1idUZaSXl4SG1KbjcweWM0OFNv?oc=5" target="_blank">13 Assessing the Quality and Reliability of Online Health Information for Breast Pathology: Search Engines Versus Large Language Models</a>&nbs
Large language models often prioritize Western moral values, overlooking other cultures - The Conversation
<a href="https://news.google.com/rss/articles/CBMilgFBVV95cUxQVmFOcTRaMmQ3Ry1zVE9YN19VZHNYVGhESTlNdVY2MkFhQXM1RXB3b2JZZ3kxZDlwZVR0aUdGcXJCMW9Fa3RRMllFVVpxRS12N3pVVmdubFBQTkJVZExFTEhXLTBrLUI2ZDJCWHN4dklHS1NGZm5ZS1oxSnp3Wml6RXNPT0pfY0t5RnZOQUVzSHdqa0VuWVE?oc=5" target="_blank">Large language models often prioritize Western moral values, overlooking other cultures</a> <font color="#6f6f6f">Digital Information World</font>
Large Language Models in Life Science Research: What Scientists Need to Know - Technology Networks
<a href="https://news.google.com/rss/articles/CBMi1AFBVV95cUxOMzJHUG0tcGFUMHR0TDdLcHozTExkU3ZTZXAwSU1ESVBlNkJGbEpvWjB5cVl1M0RqRklnYWVFY2pvMThLdWUyN0xlNXZJcG4zT1RCRnlPZWJaRUlFN0x1b3lhMXp4U1JTZzJfMk9PR3BqX21ZYlN5cVZYVXRuUmlMN0JyUUJIYndiY1NtVFNyMEJ2RGY1alRLaFNFZDhDY3IxWS04djVHcC1odkgzdlZwY3YxZFlNejBMTzV1SWFRLWxkc3pmeXItenRkTTltWG56QUYzSw?oc=5" target="_blank">Large Language Models in Life Science Research: What Scientists Need to Know</a> <font color="#6f6f6f">Technology Networks</font>
