中文 Source Card

RL训练中为什么熵减往往意味着训练收敛?

skydownacai · Sep 14, 2025, 5:09 a.m.

Sep 14, 2025, 5:09 a.m.·skydownacaiscore 68.3post-trainingpost-trainingagentsmodels

RL训练中为什么熵减往往意味着训练收敛?

最近半年以来,有关于RL+Entropy的研究非常的多。对于离散的动作空间 \mathcal{A} , 策略 \pi 在状态 s 处的entropy为 \mathcal{H} \left( \pi \left( \cdot |s \right) \right) :=\mathbb{E} _{a\sim \pi \left( \cdot |s \right)}\left[ -\log \pi \left(…

English triage

Chinese technical note: RLtraining中Why熵减往往意味着training收敛?

skydownacai published a Zhihu article relevant to post-training and RL. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.

Zhihuarticle4 upvotes

原文链接

← English feed · 中文瀑布流