Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · yesterday

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a var

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 48%RL without TD learning
  • FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail

    Shared author/contributor keys: cheng

  • LinkedLinked via arxiv author · 85%Zijie Cheng

    A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Xiang Li

    A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Hanyang Peng

    A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Zhihua Zhang

    A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics