repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · yesterday
modular/modular
The Modular Platform (includes MAX & Mojo)
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Mojo Programming Language to Transition to Open-Source Model by ModCon '26 - HackerNoon →
- FuzzyOverlapping authors or contributors · 62%Experience Memory Graph: One-Shot Error Correction for Agents →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%AIMO Interpretability Challenge →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Scalable Visual Pretraining for Language Intelligence →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%NodeImport: Imbalanced Node Classification with Node Importance Assessment →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs →
“Shared author/contributor keys: liu”
Covers
Implements
paperExperience Memory Graph: One-Shot Error Correction for AgentspaperAIMO Interpretability ChallengepaperKaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space CorrelationspaperScalable Visual Pretraining for Language IntelligencepaperNodeImport: Imbalanced Node Classification with Node Importance AssessmentpaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperGroc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMspaperM$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long StreamingpaperMusic-to-Dance Generation via Atomic MovementspaperAspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency RegularizationpaperRainDancer: RGB-Event Video Deraining with Rain-Oriented Spiking DynamicspaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperUnleashing Multimodal Large Language Models for Training-free HOI Detection in the WildpaperDKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured DatapaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperThe Seriality Gap in Video Diffusion ModelspaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperText-Driven 3D Indoor Scene Synthesis in Non-Manhattan EnvironmentspaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperBeyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse ParsingpaperTrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI SystemspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperHigh-dimensional Embedding Prior for Noisy K-space Domain MRIReconstructionpaperRFMSR: Residual Flow Matching for Image Super-ResolutionpaperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperDiffusion-GR2: Diffusion Generative Reasoning Re-rankerpaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperDisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and EditingpaperPhysMani: Physics-principled 3D World Model for Dynamic Object ManipulationpaperInside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiaspaperReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningpaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperWILDTRACE: Benchmarking Natural Evidence Trails in Long-Context ReasoningpaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMspaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperDemoPSD: Disagreement-Modulated Policy Self-DistillationpaperDecompRL: Solving Harder Problems by Learning Modular Code GenerationpaperPeak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic AssessmentpaperStable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory GatespaperDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety GuardrailpaperWavelet-Guided Semantic Signal Compensation for Inversion-Free Image EditingpaperAn Experimental Design Approach to Evaluating Agentic AI's Autonomous Model DiscoverypaperQuantum Spectral Anomaly DetectionpaperTowards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningpaperLongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction SynthesispaperELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and GenerationpaperDynamic Neural Graph Encoding of Inference Processes in Deep Weight SpacepaperSpider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL WorkflowspaperUnsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site DatasetspaperUniVR: Thinking in Visual Space for Unified Visual ReasoningpaperReal-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and NavigationpaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperExtreme Adaptive Transformer for Time Series ForecastingpaperDecoupling Language Guidance from Backbones for Text-Guided Medical SegmentationpaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperQuantum vs. Classical Machine Learning: A Unified Empirical ComparisonpaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingpaperReproducing human biases in route choice using large language models: Toward scalable behavioral modelingpaperPoint Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)paperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperMetacognition in LLMs: Foundations, Progress, and OpportunitiespaperZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank SubspacespaperDeep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning ModelspaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperLongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetpaperAlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation LearningpaperVideo = World + Event StreampaperFrequency-Structured Field Learning for Light-Field Disparity EstimationpaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperCRISP: Constrained Refinement via Iterative Squeezing Process for Robust Medical Image Segmentation under Domain ShiftpaperSciDiagramEdit: Learning to Edit Scientific Diagrams from Paper RevisionspaperHoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven ReasoningpaperStructural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-IdentificationpaperWebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web SearchpaperLift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware ManipulationpaperUI2App: Benchmarking Visual Interaction Inference in Executable Web Application GenerationpaperPACE: A Proxy for Agentic Capability EvaluationpaperTowards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment GuidancepaperNative Video-Action Pretraining for Generalizable Robot ControlpaperWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemorypaperLCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target DetectionpaperQuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response CalibrationpaperAn Exam for Active ObserverspaperSciForge: An AI-Native, Multimodal Workbench for Scientific DiscoverypaperDELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model EmbeddingspaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspaperFVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video GenerationpaperToward Semantic Communication for Real-time Mobile 3D ReconstructionpaperAudio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex VideospaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperStraight-Path Flow Matching for Incomplete Multi-View ClusteringpaperAPRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern ParadigmspaperFourier Geometric Wind Power Forecasting with Numerical Weather PredictionpaperToward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference AlignmentpaperHow Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak ContributionspaperFenced Citation-Context Retrieval for Case Law: Temporal Leakage and Degree Control Across Two JurisdictionspaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking TrainingpaperTrace-Based On-Policy Distillation for Masked Diffusion Language ModelspaperCross-Coordinate Correspondence Pruning for Image-to-Point Cloud RegistrationpaperSTBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMspaperWhen Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video GenerationpaperHarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction SynthesispaperFlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicrypaperWhat Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure ClassificationpaperAlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learningpaperThree-Body Scattering for Generative ModelingpaperO-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and ReasoningpaperSimple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspaperColor Pass-Through via Camera-Display CouplingpaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpaperUR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress ProxiespaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperMeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party MeetingspaperDAIS: Dependency-Aware Intermediate QA Supervision for Complex ReasoningpaperComputational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and ChallengespaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperISO: An RLVR-Native Optimization StackpaperStaypoint Detection from Noisy Trajectory Data [Experiment Paper]paperOmniReasoner: Thinking with Long Audio-Video via Native Tool UsepaperNo Training, Better Flights: Test-Time Scaled VLMs for UAV NavigationpaperIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerpaperABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUpaperText Template Tokens Are Implicit Semantic Registers in Diffusion TransformerspaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperXALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha DiscoverypaperGenerative AI floods and dilutes the market for bookspaperAudio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningpaperStreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video GenerationpaperNotes to Self: Can LLMs Benefit from Experiential Abstractions?paperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperPhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary MonitoringpaperOLEDLM: A Unified Language Model for OLED Molecular DesignpaperSelf Gradient Forcing: Native Long Video ExtrapolationpaperDiverse-Intent Multi-Turn Fashion Image RetrievalpaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperVera: Identity-Faithful Human Subject-to-Video GenerationpaperPerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous DrivingpaperGaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy PredictionpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperGeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process SupervisionpaperROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative ModellingpaperFrom Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIspaperMean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score EntropypaperElasticTTT: Prior-Preserving Test-Time Tuning for Video EditingpaperSPDCN: Strip-based Deformable Convolutional Network for Steel Surface Defect SegmentationpaperASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy SegmentationpaperMemTools: A Unified Research Framework for Interoperable Agent Memorypaper3D-Aware VLMs with Implicit and Explicit GeometriespaperBeyond Sufficiency: Time Series Explanation with Counterfactual NecessitypaperZero-Flow Two-Sample TestspaperSANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationpaperMedGame: Storytelling Gamification Empowered by Large Language Models for Medical EducationpaperLeveraging Extragradient for Effective Sharpness-Aware Minimization in Deep LearningpaperSAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation VerbalizationpaperMeasuring Task-Agnostic Training Data Influence Across Language Model PretrainingpaperHumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking BenchmarkpaperIntervention-Aware Clinical World Model for Post-Op Outcome Forecasting in CardiologypaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZpaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
contributed_to (incoming)
personabduldpersonJoeLoserpersonlattnerpersonKCaverlypersonatomicapple0personkeithpersondukebwpersonmodularbotpersonlaszlokindratpersonAhajhapersonhengjiewpersonConnorGraypersontboerstadpersonscottamainpersonAustinDoolittlepersonjackospersontzhenghaopersonlshpersonk-w-wpersonitramblepersonbethebunnypersonakirchhoff-modularpersonraiseirqlpersonjoshpetersonpersonjiex-liu
Related across the graph
paperDKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured DatapersonJoeLoserpaperScalable Visual Pretraining for Language IntelligencepaperToward Semantic Communication for Real-time Mobile 3D ReconstructionpaperText-Driven 3D Indoor Scene Synthesis in Non-Manhattan EnvironmentspaperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionpaperAlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learningpaperChatImage: Navigating Long-Form LLM Answers through Interactive ImagespaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperBeyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse ParsingpaperOmniReasoner: Thinking with Long Audio-Video via Native Tool UsepaperHow Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak ContributionspaperAlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation LearningpaperHow Do VLMs Fail? Vision-Operation Misalignment in Compositional VQApaperTrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI SystemspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspersonbethebunnypersonKCaverlypaperHigh-dimensional Embedding Prior for Noisy K-space Domain MRIReconstructionpaperDiffusion-GR2: Diffusion Generative Reasoning Re-rankerpaperGaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy PredictionpaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperAspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency RegularizationpaperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperMedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical ConsultationpaperDisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and EditingpaperPhysMani: Physics-principled 3D World Model for Dynamic Object ManipulationpaperISO: An RLVR-Native Optimization StackpaperAn Exam for Active ObserverspaperRFMSR: Residual Flow Matching for Image Super-ResolutionpaperInside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiaspaperReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningpaperKnowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMspersonitramblepaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperWILDTRACE: Benchmarking Natural Evidence Trails in Long-Context ReasoningpersonhengjiewpaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperPerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous DrivingpaperHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMspaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperStructural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-IdentificationpaperFVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video GenerationpaperPeak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic AssessmentpaperAudio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex VideospaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperFourier Geometric Wind Power Forecasting with Numerical Weather PredictionpaperThe Seriality Gap in Video Diffusion ModelspaperDemoPSD: Disagreement-Modulated Policy Self-DistillationpaperOLEDLM: A Unified Language Model for OLED Molecular DesignpaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspersontzhenghaopaperHoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven ReasoningpaperVera: Identity-Faithful Human Subject-to-Video GenerationpersonjackospaperStable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory GatespaperGroc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMspaperMeasuring Task-Agnostic Training Data Influence Across Language Model PretrainingpaperNodeImport: Imbalanced Node Classification with Node Importance AssessmentpaperDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety GuardrailpaperO-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and ReasoningpaperWavelet-Guided Semantic Signal Compensation for Inversion-Free Image EditingpersonlaszlokindratpaperFlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicrypersonAhajhapaperAudio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningpaperSPDCN: Strip-based Deformable Convolutional Network for Steel Surface Defect SegmentationpaperAn Experimental Design Approach to Evaluating Agentic AI's Autonomous Model DiscoverypaperQuantum Spectral Anomaly DetectionpaperLongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction SynthesispaperSANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationpaperTowards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningpaperELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and GenerationpaperSimple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMspaperDynamic Neural Graph Encoding of Inference Processes in Deep Weight SpacepaperAVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video CaptioningpaperUnsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site DatasetspaperExtreme Adaptive Transformer for Time Series ForecastingpaperText Template Tokens Are Implicit Semantic Registers in Diffusion TransformerspaperNative Video-Action Pretraining for Generalizable Robot ControlpaperAPRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern ParadigmspaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperUniVR: Thinking in Visual Space for Unified Visual ReasoningpaperDecompRL: Solving Harder Problems by Learning Modular Code GenerationpaperReal-Time Visual Intelligence on Low-Cost UAVs: A Modular Approach for Tracking, Scanning, and NavigationpaperSciForge: An AI-Native, Multimodal Workbench for Scientific DiscoverypaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperSAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation VerbalizationpaperKaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space CorrelationspaperPhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary MonitoringpersonabduldpaperSpider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL WorkflowspaperZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank SubspacespaperWebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web SearchpaperTowards Robustness against Typographic Attack with Training-free Concept LocalizationpaperMetacognition in LLMs: Foundations, Progress, and OpportunitiespaperUR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress ProxiespaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpersonlattnerpaperLift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware ManipulationpaperNo Training, Better Flights: Test-Time Scaled VLMs for UAV NavigationpersondukebwpersonjoshpetersonpaperWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemorypaperUI2App: Benchmarking Visual Interaction Inference in Executable Web Application GenerationpaperMoHallBench: A Benchmark for Motion Hallucination in Video Large Language ModelspaperQuantum vs. Classical Machine Learning: A Unified Empirical Comparisonpersonakirchhoff-modularpaperAIMO Interpretability ChallengepersonkeithpaperUltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingpaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspaperDeep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning ModelspaperFrequency-Structured Field Learning for Light-Field Disparity EstimationpaperReproducing human biases in route choice using large language models: Toward scalable behavioral modelingpaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperFenced Citation-Context Retrieval for Case Law: Temporal Leakage and Degree Control Across Two JurisdictionspaperSelf Gradient Forcing: Native Long Video ExtrapolationpaperPACE: A Proxy for Agentic Capability EvaluationpaperLeveraging Extragradient for Effective Sharpness-Aware Minimization in Deep LearningpaperVideo = World + Event StreampaperTowards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment GuidancenewsMojo Programming Language to Transition to Open-Source Model by ModCon '26 - HackerNoonpaperUnleashing Multimodal Large Language Models for Training-free HOI Detection in the WildpaperWhen Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video GenerationpaperROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative ModellingpaperDAIS: Dependency-Aware Intermediate QA Supervision for Complex ReasoningpaperCRISP: Constrained Refinement via Iterative Squeezing Process for Robust Medical Image Segmentation under Domain ShiftpaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperMusic-to-Dance Generation via Atomic MovementspaperXALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha DiscoverypaperM$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long StreamingpaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperElasticTTT: Prior-Preserving Test-Time Tuning for Video EditingpaperColor Pass-Through via Camera-Display CouplingpaperHow Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionpaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUpaperDiverse-Intent Multi-Turn Fashion Image RetrievalpaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperGeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process SupervisionpersonAustinDoolittlepaperIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerpaperStraight-Path Flow Matching for Incomplete Multi-View ClusteringpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperDeep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical TranslationpersonmodularbotpaperQuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response CalibrationpaperLCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target DetectionpaperMean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score EntropypaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspersonConnorGraypaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperZero-Flow Two-Sample TestspaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperDecoupling Language Guidance from Backbones for Text-Guided Medical SegmentationpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperSTBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMspaperNotes to Self: Can LLMs Benefit from Experiential Abstractions?paperBeyond Sufficiency: Time Series Explanation with Counterfactual Necessitypersonatomicapple0persontboerstadpaperRainDancer: RGB-Event Video Deraining with Rain-Oriented Spiking DynamicspaperPoint Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)paperASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy SegmentationpaperMemTools: A Unified Research Framework for Interoperable Agent MemorypaperCross-Coordinate Correspondence Pruning for Image-to-Point Cloud RegistrationpaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking TrainingpaperWhat Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure ClassificationpaperEdit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZpaperMeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetingspersonraiseirqlpersonjiex-liupersonlshpersonscottamainpaperIntervention-Aware Clinical World Model for Post-Op Outcome Forecasting in CardiologypaperDELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model EmbeddingspaperFrom Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIspaperLongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetpaperToward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference AlignmentpaperHarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction SynthesispaperComputational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and ChallengespaperStreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video GenerationpaperTrace-Based On-Policy Distillation for Masked Diffusion Language ModelspaperThree-Body Scattering for Generative ModelingpaperStaypoint Detection from Noisy Trajectory Data [Experiment Paper]paperHumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking BenchmarkpaperGenerative AI floods and dilutes the market for bookspersonk-w-wpaper3D-Aware VLMs with Implicit and Explicit GeometriespaperSciDiagramEdit: Learning to Edit Scientific Diagrams from Paper RevisionspaperMedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
