newsGoogle News — Machine LearningTrust 62 · AggregatorPublished 1mo agoLive · 1mo ago
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization - Apple Machine Learning Research
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Apple Machine Learning Research
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%A Self-Supervised Learning Framework for Video Encoding Complexity Clustering →
- PossiblePossibly related (embedding) · 49%MLVC: Multi-platform Learned Video Codec for Real-World Deployment →
- PossiblePossibly related (embedding) · 49%Enhanced Neural Video Representation Compression across Extreme Complexity and Quality Scales →
- PossiblePossibly related (embedding) · 48%LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models →
- PossiblePossibly related (embedding) · 48%Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection →
- PossiblePossibly related (embedding) · 51%When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression →
- PossiblePossibly related (embedding) · 46%LDJ-creat/video-helper →
- PossiblePossibly related (embedding) · 49%microsoft/SwiftStreamingMarkdown →
Covers
paperA Self-Supervised Learning Framework for Video Encoding Complexity ClusteringpaperMLVC: Multi-platform Learned Video Codec for Real-World DeploymentpaperEnhanced Neural Video Representation Compression across Extreme Complexity and Quality ScalespaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language ModelspaperDataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection
Covers (incoming)
paperWhen Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High CompressionrepoLDJ-creat/video-helperrepomicrosoft/SwiftStreamingMarkdownpaperDo All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language RetrievalpaperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperFlowMark: Mask-Guided Video WatermarkingpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionrepoSamurAIGPT/Text-To-Video-AIpaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperHOMIE: Human-object Centric Video Personalization via Multimodal Intelligent EnchancementrepoSamurAIGPT/AI-Youtube-Shorts-Generator
Related across the graph
paperG2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementpaperLongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Modelsrepomicrosoft/SwiftStreamingMarkdownpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformersrepoSamurAIGPT/AI-Youtube-Shorts-GeneratorpaperWhen Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High CompressionpaperHOMIE: Human-object Centric Video Personalization via Multimodal Intelligent EnchancementpaperMLVC: Multi-platform Learned Video Codec for Real-World DeploymentpaperA Self-Supervised Learning Framework for Video Encoding Complexity ClusteringpaperFlowMark: Mask-Guided Video WatermarkingrepoSamurAIGPT/Text-To-Video-AIpaperVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingpaperDo All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language RetrievalpaperEnhanced Neural Video Representation Compression across Extreme Complexity and Quality ScalesrepoLDJ-creat/video-helperpaperDataset Biases and Shortcut Learning in Motion-Based AI-Generated Video DetectionpaperVisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
