Slicing Return Distributions: A Practical Route to Multivariate Distributional RL
Sliced divergences make multivariate return distributions tractable by reducing them to one-dimensional projections, but TD learning still depends on how the Bellman target is sampled from the environment.