A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a var
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%RL without TD learning →
- FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail →
“Shared author/contributor keys: cheng”
- LinkedLinked via arxiv author · 85%Zijie Cheng →
“A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Xiang Li →
“A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Hanyang Peng →
“A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Zhihua Zhang →
“A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning”
