Read original ↗
paperarXivTrust 82 · PrimaryPublished 9d agoLive · 9d ago

Disclosure-Gated User Simulation for Companion-Agent Evaluation

Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learne

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “Disclosure-Gated User Simulation for Companion-Agent Evaluat” ≈ “AgentCore-8B”

  • LinkedLinked via arxiv author · 85%Yiyao Liu

    Disclosure-Gated User Simulation for Companion-Agent Evaluation

  • LinkedLinked via arxiv author · 85%Yu He

    Disclosure-Gated User Simulation for Companion-Agent Evaluation

  • FuzzySimilar title/name (fuzzy) · 59%iflytek/astron-agent

    Fuzzy title match (0.73): “Disclosure-Gated User Simulation for Companion-Agent Evaluat” ≈ “iflytek/astron-agent”

  • FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent

    Fuzzy title match (0.94): “Disclosure-Gated User Simulation for Companion-Agent Evaluat” ≈ “SWE-agent/SWE-agent”

  • FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent

    Fuzzy title match (0.94): “Disclosure-Gated User Simulation for Companion-Agent Evaluat” ≈ “zhayujie/CowAgent”

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzySimilar title/name (fuzzy) · 59%NousResearch/hermes-agent

    Fuzzy title match (0.73): “Disclosure-Gated User Simulation for Companion-Agent Evaluat” ≈ “NousResearch/hermes-agent”

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics