repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 22d ago
Zeyi-Lin/HivisionIDPhotos
⚡️HivisionIDPhotos: a lightweight and efficient AI ID photos tools. 一个轻量级的AI证件照制作算法。
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 45%Better Images of AI →
- FuzzyOverlapping authors or contributors · 62%From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Scalable Visual Pretraining for Language Intelligence →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%PanoWorld: Real-World Panoramic Generation →
“Shared author/contributor keys: lin”
Covers
Implements
paperFrom RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image ModelspaperTowards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language ModelspaperScalable Visual Pretraining for Language IntelligencepaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperPost-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT CalibrationpaperPanoWorld: Real-World Panoramic GenerationpaperDynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented GenerationpaperGatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series ForecastingpaperWat3R: Underwater 3D Geometry Learning without AnnotationspaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperSynthetic-to-Real Translation for Class-Agnostic Motion PredictionpaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative DecodingpaperReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D ReconstructionpaperStable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory GatespaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperCan LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment ReproductionpaperSIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement LearningpaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperScaling Behavior Foundation Model for Humanoid RobotspaperBeyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial UnderstandingpaperWeakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based DiffusionpaperVideo = World + Event StreampaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart GenerationpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperRethinking Quantum Continual Learning with Quantum Fisher InformationpaperDepthART: Scaling Foundation Monocular Depth to Tiny ModelspaperCross-Coordinate Correspondence Pruning for Image-to-Point Cloud RegistrationpaperAutoregressive B-Rep Shape Generation with Parametric SurfacespaperSparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionpaperThree-Body Scattering for Generative ModelingpaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperWhat Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure ClassificationpaperMulti-Modal, Multi-Environment Machine Teaching for Robust Reward LearningpaperMemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon ConversationspaperCR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene RetrievalpaperGATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape RetrievalpaperWavefront Parallelization for Efficient Learned Image CompressionpaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerpaperFlash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier TransformpaperTexture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion ModelpaperMedGame: Storytelling Gamification Empowered by Large Language Models for Medical EducationpaperLocalize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and EditspaperHumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking BenchmarkpaperSNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation
Covers (incoming)
Related across the graph
paperFrom RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image ModelspaperScalable Visual Pretraining for Language IntelligencepaperPanoWorld: Real-World Panoramic GenerationnewsAutomatically redact PII in images with Amazon NovapaperDynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented GenerationpaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperGatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series ForecastingpaperWat3R: Underwater 3D Geometry Learning without AnnotationspaperTexture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion ModelpaperWavefront Parallelization for Efficient Learned Image CompressionpaperLocalize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and EditsnewsReal-time dental image verification with Amazon SageMaker AI at Henry Schein OnepaperMonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM AdaptationpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperSynthetic-to-Real Translation for Class-Agnostic Motion PredictionpaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative DecodingpaperPost-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT CalibrationpaperMM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationpaperStable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory GatespaperReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D ReconstructionnewsArtificial Intelligence (AI) Otoscopic Image Triaging - openPR.compaperFlash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier TransformpaperCan LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment ReproductionpaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperMulti-Modal, Multi-Environment Machine Teaching for Robust Reward LearningpaperWeakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based DiffusionpaperSIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement LearningpaperMemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon ConversationspaperOn Success and Simplicity: A Second Look at Transferable Vision-Language Attack PipelinepaperVideo = World + Event StreampaperVGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy PredictionpaperScaling Behavior Foundation Model for Humanoid RobotspaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperGATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape RetrievalpaperIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerpaperEAGLE-360: Embodied Active Global-to-Local Exploration in 360$^\circ$paperChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart GenerationpaperRethinking Quantum Continual Learning with Quantum Fisher InformationpaperSparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionpaperBeyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial UnderstandingpaperCross-Coordinate Correspondence Pruning for Image-to-Point Cloud RegistrationpaperTowards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language ModelsnewsBetter Images of AIpaperWhat Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure ClassificationpaperCR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene RetrievalpaperAutoregressive B-Rep Shape Generation with Parametric SurfacespaperDepthART: Scaling Foundation Monocular Depth to Tiny ModelspaperThree-Body Scattering for Generative ModelingpaperHumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking BenchmarkpaperSNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame InterpolationpaperMedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
