Deep Reinforcement Learning at the Edge of the Statistical Precipice

deep-rlevaluationstatistical-analysisbenchmarkingatarirliable

Abstraction: Critique of deep RL evaluation using point estimates; proposes robust statistical methodology

Key points:

Connections: Rliable · Reinforcement Learning · Statistical Evaluation · Benchmarking

Source: https://arxiv.org/abs/2108.13264