中文 Source Card

LLM Agent RL的一些实践感悟

skydownacai · Dec 4, 2025, 2:15 a.m.

Dec 4, 2025, 2:15 a.m.·skydownacaiscore 68.3post-trainingagentsmodelspost-training

LLM Agent RL的一些实践感悟

最近几个月在各种场景上做了大量的agent RL训练,例如search agent, 数据分析agent 等等。既有小的dense model,也有大的moe。有单一数据场景,有多源合板数据场景。有失败的经历,也有成功的经历。这里抽空分享下一些感悟,比较乱,随便看看就行: 稳定性…

English triage

Chinese technical note: LLM agents RL的一些实践感悟

skydownacai published a Zhihu article relevant to post-training and RL, AI agents and coding workflows, frontier and open model development. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.

Zhihuarticle23 upvotes

原文链接

← English feed · 中文瀑布流