repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 22d ago
bytedance/deer-flow
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%Experience Memory Graph: One-Shot Error Correction for Agents →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0 →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents →
“Shared author/contributor keys: wang”
Implements
paperWan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance GenerationpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperDo Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0paperSimon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-ResolutionpaperSelf-Evolving Agent Harnesses via Gated Semantic Quality-DiversitypaperFrom RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image ModelspaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous EnvironmentspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperM$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long StreamingpaperAspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency RegularizationpaperTerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at ScalepaperLearning Mechanistic Reasoning for Chemical Reactions with Large Language ModelspaperScalable Visual Pretraining for Language IntelligencepaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperBeyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph AnalysispaperAn Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and GenerationpaperDKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured DatapaperCan Induced Emotion Bias LLM Behaviors in Sequential Decision Making?paperMachine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based MethodologypaperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperText-Driven 3D Indoor Scene Synthesis in Non-Manhattan EnvironmentspaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Modelspaper4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene PerceptionpaperSecure Decentralized Federated Learning via Gossip and Virtual VotingpaperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent SystemspaperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperCortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon ManipulationpaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperDisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and EditingpaperPhysMani: Physics-principled 3D World Model for Dynamic Object ManipulationpaperSynthetic-to-Real Translation for Class-Agnostic Motion PredictionpaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperInside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiaspaperAir Quality Downscaling with Station-Guided Pseudo-SupervisionpaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperGaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TaskspaperAccelerated Mixing Time of Randomized Hamiltonian Monte CarlopaperFinding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy ReportspaperPurified OPSD: On-Policy Self-Distillation Without Losing How to ThinkpaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperOptimizing Visual Generative Models via Distribution-wise RewardspaperCurateEvo: Data-Curation Evolving for Agentic Post-TrainingpaperReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D ReconstructionpaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperIdeas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationpaperDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety GuardrailpaperTowards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningpaperDynamic Neural Graph Encoding of Inference Processes in Deep Weight SpacepaperUniVR: Thinking in Visual Space for Unified Visual ReasoningpaperControllable Sim Agents with Behavior LatentspaperWhen Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?paperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperCausalMix: Data Mixture as Causal Inference for Language Model TrainingpaperFADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face RestorationpaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingpaperWhat Images Cannot Say: Language-Guided Olfactory Representation LearningpaperLean-QIT: Towards a Formal Infrastructure for Quantum Information TheorypaperDemonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at ScalepaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperUnderstanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID DatapaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionpaperPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel SpacepaperZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank SubspacespaperMIRAGE: Defending Long-Form RAG Against Misinformation PollutionpaperScaling Behavior Foundation Model for Humanoid RobotspaperBrainPilot: Automating Brain Discovery with Agentic ResearchpaperLQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics ResearchpaperRubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise EvidencepaperBadWAM: When World-Action Models Dream Right but Act WrongpaperBenchmarking Multimodal Large Language Models for Scientific Visualization LiteracypaperDAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image ClassificationpaperRoGS: Adaptive Meshgrid Gaussian for Large-Scale Road Surface MappingpaperVideo = World + Event StreampaperJADE-GS: Joint Alternating Deblurring Guided by Events in 3D Gaussian SplattingpaperFrom Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast PlantingpaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperSciDiagramEdit: Learning to Edit Scientific Diagrams from Paper RevisionspaperMotion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from EchocardiographypaperStructural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-IdentificationpaperWhen Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk SpacepaperDetecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought AuditingpaperTowards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment GuidancepaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemorypaperWebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web SearchpaperCycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle ConsistencypaperAn Exam for Active ObserverspaperContextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech ComprehensionpaperBetter Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed ReasoningpaperAn MLIR-Based Compilation Method for Large Language ModelspaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpaperJoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA ModelspaperCRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning DatapaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperDPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense PredictionpaperBeyond Unfolding: 60x Faster One-Stage Unmixing for Closely-Spaced Infrared Small TargetspaperStraight-Path Flow Matching for Incomplete Multi-View ClusteringpaperGeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector TrainingpaperGeometry and Gradient-based Partitioning for Panoramic Outdoor ReconstructionpaperPaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper RoutingpaperActive rejection enables reliable generalization of universal machine-learning interatomic potentialspaperLenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce LearningpaperDistilled Reinforcement Learning for LLM Post-trainingpaperNoise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label OptimizationpaperDepthART: Scaling Foundation Monocular Depth to Tiny ModelspaperToward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference AlignmentpaperHow Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak ContributionspaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking TrainingpaperCross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image EditingpaperSTBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMspaperAdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature SuppressionpaperALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable EnvironmentspaperAutoregressive B-Rep Shape Generation with Parametric SurfacespaperSparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionpaperSWE-Pruner Pro: The Coder LLM Already Knows What to PrunepaperSelectInfer: Selective Neuron Loading and Computation for On-Device LLMspaperGigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment AnalysispaperEnhancing Rubric-based RL via Self-DistillationpaperEVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain DatabasepaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspaperThe Many Senses of Visual Similarity: A Text-Prompted Image Perceptual MetricpaperLossless-INR: Lossless Volumetric Implicit Neural RepresentationspaperO-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and ReasoningpaperOcclusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level AttentionpaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperUniETP: Unifying Environments for Generalizable Embodied Task PlanningpaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingpaperAutoMem: Automated Learning of Memory as a Cognitive SkillpaperMBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation ModelspaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperVaseMuseum: Digital Intelligent Museum for Ancient Greek PotterypaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperSequential Learner Modeling Using Multi-Relational Graph Convolutional NetworkspaperCopy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement LearningpaperMeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party MeetingspaperDAIS: Dependency-Aware Intermediate QA Supervision for Complex ReasoningpaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperISO: An RLVR-Native Optimization StackpaperOmniReasoner: Thinking with Long Audio-Video via Native Tool UsepaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperNo Training, Better Flights: Test-Time Scaled VLMs for UAV NavigationpaperCR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene RetrievalpaperGATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape RetrievalpaperBeyond the Single Camera: Agentic Multi-View Reasoning in Sports Video UnderstandingpaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperThe Geometry of Memorization: Finite-Time Spectral Sensitivity as a Diagnostic for Flow Matching ModelspaperCAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal ConsistencypaperXALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha DiscoverypaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperAudio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningpaperOn the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural LenspaperLabel-Free Finite-Volume-Residual Training of Attention Graph Neural Networks for Coupled Thermo-Fluid FieldspaperPhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary MonitoringpaperPercepCap: Video Captioner with Structured Spatio-Temporal PerceptionpaperOLEDLM: A Unified Language Model for OLED Molecular DesignpaperLook Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMspaperEvolving Cache Schedules for Fast Diffusion Policy InferencepaperVera: Identity-Faithful Human Subject-to-Video GenerationpaperRS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image EditingpaperPerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous DrivingpaperRIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVspaperGaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy PredictionpaperPushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching RenderingpaperExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian SplattingpaperAlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learningpaperGeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process SupervisionpaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperAREX: Towards a Recursively Self-Improving Agent for Deep ResearchpaperPATS: Policy-Aware Training Scaffolding for Agentic Reinforcement LearningpaperAdaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMspaperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciespaperToward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language ModelspaperTexture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion ModelpaperTowards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert ApproachpaperASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy SegmentationpaperFlash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier TransformpaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperBeyond Sycophancy: Structured Resistance and Compliance in LLM Moral ReasoningpaperZero-Flow Two-Sample Tests
contributed_to (incoming)
personMagicCubepersonhetaoBackendpersonWillemJiangpersonhenry-bytedpersonLofiSupersondependabot[bot]personforelevenpersonfancyboi999personggnnggezpersonShenAC-SACpersonLittleChenLiyapersonHuixin615personEilen6316personsyzhang622persongreatmengqipersonly-wang19personyangzhelipersonforx11person18062706139fczpersonxunliupersonJasonOA888personGujiasshpersonknuknYpersonocto-patchpersonwhhe
Covers (incoming)
Related across the graph
paperSimon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-ResolutionpersonJasonOA888paperWan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance GenerationpaperFrom RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image ModelspaperDKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured DatapaperTerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at ScalepaperScalable Visual Pretraining for Language IntelligencepersonforelevenpaperText-Driven 3D Indoor Scene Synthesis in Non-Manhattan EnvironmentspaperLabel-Free Finite-Volume-Residual Training of Attention Graph Neural Networks for Coupled Thermo-Fluid FieldspaperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperAlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learningpaperSelectInfer: Selective Neuron Loading and Computation for On-Device LLMspaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperOmniReasoner: Thinking with Long Audio-Video via Native Tool UsepaperHow Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak ContributionspaperSPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Modelspaper4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene PerceptionpaperSecure Decentralized Federated Learning via Gossip and Virtual VotingpaperThe Many Senses of Visual Similarity: A Text-Prompted Image Perceptual MetricpaperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent SystemspaperMachine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based MethodologypaperNoise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label OptimizationpaperCortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon ManipulationpaperGaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy PredictionpaperAspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency RegularizationpaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperCan Induced Emotion Bias LLM Behaviors in Sequential Decision Making?paperDisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and EditingpersonknuknYpaperPhysMani: Physics-principled 3D World Model for Dynamic Object ManipulationpaperISO: An RLVR-Native Optimization StackpaperSelf-Evolving Agent Harnesses via Gated Semantic Quality-DiversitypaperTowards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert ApproachpaperTexture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion ModelpaperEvolving Cache Schedules for Fast Diffusion Policy InferencepaperCross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image EditingpaperAn Exam for Active ObserverspaperBetter Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed ReasoningpaperInside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiaspaperAir Quality Downscaling with Station-Guided Pseudo-SupervisionpaperMotion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from EchocardiographypaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World Taskspersonsyzhang622paperLearning Mechanistic Reasoning for Chemical Reactions with Large Language ModelspaperCoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language ModelspaperGaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TaskspaperBrainPilot: Automating Brain Discovery with Agentic ResearchpaperContextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech ComprehensionpaperAccelerated Mixing Time of Randomized Hamiltonian Monte CarlopaperPerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous DrivingpaperAn Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and GenerationpaperStructural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-IdentificationpaperFinding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy ReportspaperSynthetic-to-Real Translation for Class-Agnostic Motion PredictionpaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpersongreatmengqipaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperOLEDLM: A Unified Language Model for OLED Molecular DesignpaperDPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense PredictionpaperBeyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph AnalysispaperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperVera: Identity-Faithful Human Subject-to-Video GenerationpaperLook Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMspaperJoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA ModelspersonyangzhelipaperOcclusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level AttentionpaperOptimizing Visual Generative Models via Distribution-wise RewardspaperCurateEvo: Data-Curation Evolving for Agentic Post-TrainingpaperReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D ReconstructionpaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperIdeas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationpaperDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety GuardrailpaperO-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and ReasoningpaperBadWAM: When World-Action Models Dream Right but Act WrongpaperAudio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningpaperFlash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier TransformpaperUniETP: Unifying Environments for Generalizable Embodied Task PlanningpaperTowards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningpaperDynamic Neural Graph Encoding of Inference Processes in Deep Weight SpacepaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingpersonhetaoBackendpaperBenchmarking Multimodal Large Language Models for Scientific Visualization LiteracypaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperEVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain DatabasepaperDistilled Reinforcement Learning for LLM Post-trainingpaperPercepCap: Video Captioner with Structured Spatio-Temporal PerceptionpaperPurified OPSD: On-Policy Self-Distillation Without Losing How to ThinkpaperUniVR: Thinking in Visual Space for Unified Visual ReasoningpaperGeometry and Gradient-based Partitioning for Panoramic Outdoor ReconstructionpaperPaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper RoutingpersonggnnggezpaperAdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature Suppressionpersonhenry-bytedpaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperCausalMix: Data Mixture as Causal Inference for Language Model TrainingpaperOn the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural LenspaperBeyond the Single Camera: Agentic Multi-View Reasoning in Sports Video UnderstandingpaperControllable Sim Agents with Behavior LatentspaperWhen Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?paperPhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary MonitoringpaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperFADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face RestorationpaperZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank SubspacespaperWebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web SearchpaperSequential Learner Modeling Using Multi-Relational Graph Convolutional NetworkspaperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperNo Training, Better Flights: Test-Time Scaled VLMs for UAV NavigationpaperUnderstanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID DatapersonHuixin615paperMBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation ModelspaperWhen Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk SpacepaperJADE-GS: Joint Alternating Deblurring Guided by Events in 3D Gaussian SplattingpaperRoGS: Adaptive Meshgrid Gaussian for Large-Scale Road Surface MappingpaperGigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment AnalysispaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemorypaperPATS: Policy-Aware Training Scaffolding for Agentic Reinforcement LearningpaperDetecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought AuditingpaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editingperson18062706139fczpaperLQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics ResearchpaperRS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image EditingpaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspersondependabot[bot]paperOne More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification PoliciespersonShenAC-SACpaperWhat Images Cannot Say: Language-Guided Olfactory Representation LearningpaperLean-QIT: Towards a Formal Infrastructure for Quantum Information TheorypaperVideo = World + Event StreampaperTowards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment GuidancepaperDemonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at ScalepaperDAIS: Dependency-Aware Intermediate QA Supervision for Complex ReasoningpaperFrom Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast PlantingpaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperScaling Behavior Foundation Model for Humanoid RobotspaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperXALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha DiscoverypaperM$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long StreamingpaperThe Geometry of Memorization: Finite-Time Spectral Sensitivity as a Diagnostic for Flow Matching ModelspaperAutoMem: Automated Learning of Memory as a Cognitive SkillpaperEnhancing Rubric-based RL via Self-DistillationpaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel SpacepaperGATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape RetrievalpaperCAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal ConsistencypaperGeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process SupervisionpaperExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesispaperBeyond Unfolding: 60x Faster One-Stage Unmixing for Closely-Spaced Infrared Small TargetspaperStraight-Path Flow Matching for Incomplete Multi-View ClusteringpaperGeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector TrainingpaperExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian SplattingpaperVaseMuseum: Digital Intelligent Museum for Ancient Greek PotterypaperActive rejection enables reliable generalization of universal machine-learning interatomic potentialspaperAREX: Towards a Recursively Self-Improving Agent for Deep ResearchpaperLenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce LearningpaperArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text RenderingpaperAn MLIR-Based Compilation Method for Large Language ModelspaperCRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning DatapaperPerceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningpersonEilen6316paperLossless-INR: Lossless Volumetric Implicit Neural RepresentationspaperCycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle ConsistencypaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperZero-Flow Two-Sample TestspaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperBeyond Sycophancy: Structured Resistance and Compliance in LLM Moral ReasoningpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperSTBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMspersonMagicCubepaperVoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversionpersonforx11paperALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable EnvironmentspaperToward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language ModelspaperSparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detectionpersonocto-patchpaperASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy SegmentationpaperCopy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement LearningpaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking TrainingpersonWillemJiangpersonGujiasshpaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApersonly-wang19paperMeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party MeetingspersonwhhenewsByteDance Enters Physical AI Arena: Unlocking the Second Wave of the Tech Industry Feast - 36 KrpaperRubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise EvidencepaperToward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference AlignmentpaperCR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene RetrievalpaperAutoregressive B-Rep Shape Generation with Parametric SurfacespaperDepthART: Scaling Foundation Monocular Depth to Tiny ModelspaperSWE-Pruner Pro: The Coder LLM Already Knows What to Prunepersonfancyboi999paperPushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching RenderingpaperRIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVspaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperDAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image ClassificationpaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspersonLofiSupaperAdaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMspersonxunliupaperDo Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0paperSciDiagramEdit: Learning to Edit Scientific Diagrams from Paper RevisionspersonLittleChenLiyapaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous EnvironmentspaperMIRAGE: Defending Long-Form RAG Against Misinformation Pollution
