Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /Unity-Technologies/ml-agents
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Unity-Technologies/ml-agents

The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 84%Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning →

    “Fuzzy title match (0.92): “Empowering GUI Agents via Autonomous Experience Exploration ” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution →

    “Fuzzy title match (0.92): “MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents v” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents →

    “Fuzzy title match (0.92): “When State Becomes an Attack Surface: State-Semantic Injecti” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding →

    “Fuzzy title match (0.92): “When Agents Coordinate: Measuring Coordination in Multi-Agen” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%Neurosymbolic Embodied Agents →

    “Fuzzy title match (0.92): “Neurosymbolic Embodied Agents” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents →

    “Fuzzy title match (0.92): “StagedWorkspace: A Versioned Workspace for Knowledge-Work Ag” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents →

    “Fuzzy title match (0.92): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “Unity-Technologies/ml-agents””

  • FuzzySimilar title/name (fuzzy) · 84%On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification →

    “Fuzzy title match (0.92): “On the Fragility of Self-Improving Agents: Variance, Task Or” ≈ “Unity-Technologies/ml-agents””

Implements

paperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperNeurosymbolic Embodied AgentspaperStagedWorkspace: A Versioned Workspace for Knowledge-Work AgentspaperPolicy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL AgentspaperOn the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationpaperAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementpaperFrom Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical DocumentationpaperBreak It Down, Pass It On: Cross-Task Skill Transfer in LLM AgentspaperThe Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and AgentspaperSpecification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software MigrationpaperCIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model AgentspaperEarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural HazardspaperSWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?paperCan Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development PipelinepaperThinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video AgentspaperCAFE: Self-Improving Search Agents Need Co-Evolving FeedbackpaperSkillForge: Evolving Verifiable Skills for Reinforcement Learning AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionpaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?paperACE: Pluggable Adaptive Context Elasticizer across AgentspaperGenerative Skill Composition for LLM AgentspaperAgents in the Wild: Where Research Meets DeploymentpaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperHarnessing Code Agents for Automatic Software VerificationpaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessespaperVero: Can AI Agents Build Formally Verified Software Repositories?paperSelf-rewarding agents that retrace failurespaperExperience Memory Graph: One-Shot Error Correction for AgentspaperDeepStress: Stress-Testing Deep Search AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous EnvironmentspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperControllable Sim Agents with Behavior LatentspaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentspaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Covers

newsAgentic AI for Robot TeamsnewsWorkato Launches Open-Source ‘Labs’ Hub for AI-Driven Automation - Open Source For You

Related to

articleA field guide to AI agents in 2026

Covers (incoming)

newsMLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AI - AiThoritynewsThe HydroGym reinforcement learning platform for fluid dynamics - NaturenewsBlueVoyant releases AI agent security service for Microsoft environments - KMWorldnewsSkyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation Benchmark That Makes Continual Reinforcement Learning Necessary Under Structured Non-Stationarity - MarkTechPost

Implements (incoming)

paperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperGaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TaskspaperDirectional Constraints for Efficient Exploration in Safe Reinforcement Learning

Related across the graph

paperOn the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationpaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperBreak It Down, Pass It On: Cross-Task Skill Transfer in LLM AgentspaperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperGaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TaskspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorynewsMLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AI - AiThoritypaperDeepStress: Stress-Testing Deep Search AgentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentsnewsSkyfall AI Releases MORPHEUS: A Persistent Enterprise Simulation Benchmark That Makes Continual Reinforcement Learning Necessary Under Structured Non-Stationarity - MarkTechPostpaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingnewsWorkato Launches Open-Source ‘Labs’ Hub for AI-Driven Automation - Open Source For YoupaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperSelf-rewarding agents that retrace failurespaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperControllable Sim Agents with Behavior LatentspaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperACE: Pluggable Adaptive Context Elasticizer across AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperSWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?paperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionpaperThe Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and AgentspaperAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementpaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperNeurosymbolic Embodied AgentspaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperGenerative Skill Composition for LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperEarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural HazardspaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentspaperHarnessing Code Agents for Automatic Software VerificationpaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperVero: Can AI Agents Build Formally Verified Software Repositories?paperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentspaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentspaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperPolicy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL AgentspaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperPlover: Steering GUI Agents through Plan-Centric InteractionnewsAgentic AI for Robot TeamspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessesarticleA field guide to AI agents in 2026paperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?paperCAFE: Self-Improving Search Agents Need Co-Evolving FeedbacknewsBlueVoyant releases AI agent security service for Microsoft environments - KMWorldnewsThe HydroGym reinforcement learning platform for fluid dynamics - NaturepaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAgents in the Wild: Where Research Meets DeploymentpaperCan Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development PipelinepaperThinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video AgentspaperFrom Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical DocumentationpaperCIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model AgentspaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperStagedWorkspace: A Versioned Workspace for Knowledge-Work AgentspaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentspaperSkillForge: Evolving Verifiable Skills for Reinforcement Learning AgentspaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperDirectional Constraints for Efficient Exploration in Safe Reinforcement LearningpaperSpecification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software MigrationpaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Knowledge path·POn the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification→PAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?→PAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents→RUnity-Technologies/ml-agents

Topics

deep-learningdeep-reinforcement-learningmachine-learningneural-networksreinforcement-learningunityunity3d

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Maintenance74
RIS82GitHub verified
Graph trust82Primary
Graph score19558