Read original ↗
paperarXivTrust 82 · PrimaryPublished 27d agoLive · 26d ago

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset, but its performance is fundamentally limited by dataset coverage. Action preference queries leverage expert feedback without additional environment interaction, enabling policy improvement during offline training. However, existing methods still face two key challenges: selecting informative preference queries and effectively exploiting the collected feedback. Current approaches typically rely only on the distance between policy actions and dataset actions for query selection, while enforcing fixed con

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 78%sgl-project/sglang

    Shared author/contributor keys: luo, zhou

  • LinkedLinked via arxiv author · 85%Li-Rong Zhou

    Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

  • LinkedLinked via arxiv author · 85%Qin-Wen Luo

    Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

  • LinkedLinked via arxiv author · 85%Sheng-Jun Huang

    Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

Implements (incoming)

authored (incoming)

Related across the graph

Topics