今日已更新 344 条资讯 | 累计 37249 条内容
关于我们

Delay-corrected Bellman operator + causal attribution for constrained RL contraction proof under unknown stochastic delay [R]

/u/No_Cauliflower7923 2026年08月24日 20:11 1 次阅读 来源:Reddit r/MachineLearning

Standard constrained RL assumes consequences are immediate and attributable to the current action. This breaks down whenever violations are delayed and stochastic, which is most real-world settings you end up penalizing whatever action happened to precede the observed violation, not the action that caused it. Working on CCPL (Causal Consequence-Penalized Learning) to address this: - A delay-corrected Bellman operator using an adaptive effective discount learned from the consequence-delay distribution. Contraction proof holds under unknown stochastic delay. - An Interventional Consequence Net (ICN), pretrained on structural-causal-model labels, estimating marginal causal contribution per action for attribution rather than penalizing based on temporal proximity. Limitations, to be upfront about them: - The ICN currently requires access to the environment's structural causal model to generate pretraining labels it's not learned end-to-end from observational or interventional data alone. That's a real constraint on applicability outside benchmark settings where the SCM is known or can be reasonably specified. Open to contributions and collaborators, especially if you work in constrained/safe RL or causal inference feel free to open an issue or reach out directly. submitted by /u/No_Cauliflower7923 [link] [留言]

本文内容来源于互联网,版权归原作者所有
查看原文