Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

NirDiamant/GenAI_Agents

50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems.

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Implements

paperSelf-rewarding agents that retrace failurespaperExperience Memory Graph: One-Shot Error Correction for AgentspaperDeepStress: Stress-Testing Deep Search AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous EnvironmentspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperHarnessing Code Agents for Automatic Software VerificationpaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessespaperVero: Can AI Agents Build Formally Verified Software Repositories?paperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionpaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentspaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?paperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingpaperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperNeurosymbolic Embodied AgentspaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperControllable Sim Agents with Behavior LatentspaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026paperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentspaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperACE: Pluggable Adaptive Context Elasticizer across AgentspaperGenerative Skill Composition for LLM AgentspaperAgents in the Wild: Where Research Meets DeploymentpaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works

Covers

Covers (incoming)

Implements (incoming)

Related across the graph

paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperAlways-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgentspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperWhen Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier ModelspaperDynamo: Dynamic Skill-Tool Evolution for Vision-Language AgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperAutoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DatapaperDanus: Orchestrating Mathematical Reasoning Agents with Fact-Graph MemorypaperDeepStress: Stress-Testing Deep Search AgentspaperLLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial ObservabilitypaperSMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingpaperMRMS: A Multi-Resolution Memory Substrate for Long-Lived AI AgentspaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperThe Ethics of Autonomous AI Agents for Offensive SecuritypaperVEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentspaperToken-Flow Firewall: Semantic Runtime Auditing for Persistent AI AgentspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperWhen Agents Lie: Premeditation, Persistence, and Exploitation in Repeated GamespaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperFlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationspaperWorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingpaperHy-Embodied-VLM-1.0: Efficient Physical-World AgentspaperSelf-rewarding agents that retrace failurespaperOpenForgeRL: Train Harness-native Agents in Any EnvironmentpaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionpaperControllable Sim Agents with Behavior LatentspaperAdvancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical AutonomypaperACE: Pluggable Adaptive Context Elasticizer across AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionnewsUnderstanding Generative AI: Beyond Chatbots and Prompts - themetropolitan.metrostate.edupaperQVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentspaperNeurosymbolic Embodied AgentspaperBioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillancepaperGenerative Skill Composition for LLM AgentspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationpaperEmpowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningpaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperBeyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentsnewsPlurality Released: fully Free and Open Source AI agents/chatbot platform for local AIpaperHarnessing Code Agents for Automatic Software VerificationpaperManimAgent: Self-Evolving Multimodal Agents for Visual EducationpaperVero: Can AI Agents Build Formally Verified Software Repositories?paperA Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code AgentsnewsUnderstanding Generative AI: Beyond Chatbots and Prompts - Metro State UniversitypaperCodeRescue: Budget-Calibrated Recovery Routing for Coding AgentspaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatespaperWhen State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied AgentspaperMemory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentspaperTask-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026newsHow Do Generative AI Tools Like ChatGPT Work? - University of Central FloridapaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentspaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperMedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation AgentspaperPlover: Steering GUI Agents through Plan-Centric InteractionpaperTowards Detecting Inconsistencies in End-to-end Generated TODspaperWhen Agents Coordinate: Measuring Coordination in Multi-Agent AI CodingpaperExperience Memory Graph: One-Shot Error Correction for AgentspaperEnhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory ProcessespaperSelf-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?paperSearching Videos as Trees: Self-Correcting Agents for Grounded Long Video QApaperAgents in the Wild: Where Research Meets DeploymentpaperEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldpaperDigital Pantheon: Simulating and Auditing Coalition Formation with LLM AgentsnewsTo Build More Believable Bots, Simulate The Neurochemistry - HackadaypaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually WorkspaperTRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentspaperOmniaBench: Benchmarking General AI Agents Across Diverse ScenariospaperDevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

Topics