Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic B
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%[2607.07508] Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning →
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Bellman Calibration for Marginalized Importance Weighting in” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Lars van der Laan →
“Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Nathan Kallus →
“Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning”
- PossiblePossibly related (embedding) · 57%Offline Reinforcement Learning Improves Through Active Model Selection and Bayesian Optimization - Bioengineer.org →
