Read original ↗repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
Unity-Technologies/ml-agents
The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
“Fuzzy title match (0.92): “DeepStress: Stress-Testing Deep Search Agents” ≈ “Unity-Technologies/ml-agents””
“Fuzzy title match (0.92): “TRACE: Turn-level Reward Assignment via Credit Estimation fo” ≈ “Unity-Technologies/ml-agents””
“Fuzzy title match (0.92): “When Does Combining Language Models Help? A Co-Failure Ceili” ≈ “Unity-Technologies/ml-agents””
“Fuzzy title match (0.92): “PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Ag” ≈ “Unity-Technologies/ml-agents””
“Fuzzy title match (0.92): “UniClawBench: A Universal Benchmark for Proactive Agents on ” ≈ “Unity-Technologies/ml-agents””
Implements
paperDeepStress: Stress-Testing Deep Search AgentspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionpaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?paperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingpaperACE: Pluggable Adaptive Context Elasticizer across AgentspaperGenerative Skill Composition for LLM AgentspaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessespaperVero: Can AI Agents Build Formally Verified Software Repositories?paperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperNeurosymbolic Embodied AgentspaperSpecification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software MigrationpaperSelf-rewarding agents that retrace failurespaperExperience Memory Graph: One-Shot Error Correction for AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous EnvironmentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperAgents in the Wild: Where Research Meets DeploymentpaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperCIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model AgentspaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperControllable Sim Agents with Behavior LatentspaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperHarnessing Code Agents for Automatic Software VerificationpaperStagedWorkspace: A Versioned Workspace for Knowledge-Work AgentspaperPolicy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL AgentspaperOn the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationpaperAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementpaperFrom Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical DocumentationpaperBreak It Down, Pass It On: Cross-Task Skill Transfer in LLM AgentspaperThe Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents Related across the graph
paperOn the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationpaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperBreak It Down, Pass It On: Cross-Task Skill Transfer in LLM AgentspaperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperGaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TaskspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorynewsMLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AI - AiThoritypaperDeepStress: Stress-Testing Deep Search AgentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentsnewsSkyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation Benchmark That Makes Continual Reinforcement Learning Necessary Under Structured Non-Stationarity - MarkTechPostpaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingnewsWorkato Launches Open-Source ‘Labs’ Hub for AI-Driven Automation - Open Source For YoupaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperSelf-rewarding agents that retrace failurespaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperControllable Sim Agents with Behavior LatentspaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperACE: Pluggable Adaptive Context Elasticizer across AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionpaperThe Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and AgentspaperAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementpaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperNeurosymbolic Embodied AgentspaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperGenerative Skill Composition for LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentspaperHarnessing Code Agents for Automatic Software VerificationpaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperVero: Can AI Agents Build Formally Verified Software Repositories?paperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentspaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperPolicy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL AgentspaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperPlover: Steering GUI Agents through Plan-Centric InteractionnewsAgentic AI for Robot TeamspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessesarticleA field guide to AI agents in 2026paperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?newsBlueVoyant releases AI agent security service for Microsoft environments - KMWorldnewsThe HydroGym reinforcement learning platform for fluid dynamics - NaturepaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAgents in the Wild: Where Research Meets DeploymentpaperFrom Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical DocumentationpaperCIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model AgentspaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperStagedWorkspace: A Versioned Workspace for Knowledge-Work AgentspaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentspaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperDirectional Constraints for Efficient Exploration in Safe Reinforcement LearningpaperSpecification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software MigrationpaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments