repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 8h ago
xlang-ai/OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 67%MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments →
- PossiblePossibly related (embedding) · 61%Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots →
- PossiblePossibly related (embedding) · 59%AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration →
- PossiblePossibly related (embedding) · 59%Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering →
- PossiblePossibly related (embedding) · 58%Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents →
- PossiblePossibly related (embedding) · 46%SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments →
- PossiblePossibly related (embedding) · 51%EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments →
- PossiblePossibly related (embedding) · 50%Multiplayer Interactive World Models with Representation Autoencoders →
Implements
paperMECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied EnvironmentspaperEmbodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous RobotspaperAirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied CollaborationpaperRehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question AnsweringpaperLearning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Implements (incoming)
paperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentspaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperMultiplayer Interactive World Models with Representation AutoencoderspaperCortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon ManipulationpaperSearch Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual GenerationpaperAlayaWorld: Long-Horizon and Playable Video World GenerationpaperTraining-Free Acceleration for Vision-Language-Action Models with Action Caching and RefinementpaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperFrom World Action Models to Embodied Brains: A Roadmap for Open-World Physical IntelligencepaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperKnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Covers (incoming)
Related across the graph
paperCortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon ManipulationpaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TasksnewsCfP | RTCA @ NeurIPS 2026 [R]paperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperAlayaWorld: Long-Horizon and Playable Video World GenerationnewsNVIDIA and Hugging Face Bring New Models and Frameworks to LeRobot for the Open Robotics CommunitypaperKnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and SkillpaperMECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied EnvironmentspaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperRehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question AnsweringpaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperLearning from Failure: Inference-Time Self-Improvement for Computer-Use AgentspaperSteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentsnewsMIRA: Multiplayer Interactive World Models trained on Rocket League [R]paperTraining-Free Acceleration for Vision-Language-Action Models with Action Caching and RefinementpaperMultiplayer Interactive World Models with Representation AutoencoderspaperEmbodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous RobotspaperAirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied CollaborationpaperFrom World Action Models to Embodied Brains: A Roadmap for Open-World Physical IntelligencepaperSearch Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
