RL训练中为什么熵减往往意味着训练收敛?
最近半年以来,有关于RL+Entropy的研究非常的多。对于离散的动作空间 \mathcal{A} , 策略 \pi 在状态 s 处的entropy为 \mathcal{H} \left( \pi \left( \cdot |s \right) \right) :=\mathbb{E} _{a\sim \pi \left( \cdot |s \right)}\left[ -\log \pi \left(…
English triage
Chinese technical note: RLtraining中Why熵减往往意味着training收敛?
skydownacai published a Zhihu article relevant to post-training and RL. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.