Blog
Notes on training language models and the systems around them, link lists, and the occasional essay. Some posts are in Chinese.
2025 5 posts
- 从 SFT 到 RLVR 的平滑过渡:GRPO 的直觉解释 这篇文章主要是从直觉角度来解释大模型训练中 SFT 和 RL 的关系。我们可以看到区别于SFT时“老师说的就是对”的学习方式,RL 是为了能够高效的利用“没那么正确”的样本,从而增加模型“正确”的可能性。这篇文章是我在阅读整理 Understanding Reinforcement Learning for Model Training, and future directions with GRAPE 这篇报告的思考和笔记,也强烈建议大家去读原文。
- A weird debugging process for Jupyter Notebook on WSL 2 事情的起因是: 在 WSL 2 下开启 Jupyter Notebook, 在 windows 下可以通过 127.0.0.1:8888 启动,但无法通过 localhost:8888 启动。
- Thoughts on vibe coding Recently, there’s a trend called “vibe coding,” proposed by Andrej Karpathy. It essentially means that by just telling what you want to build to LLMs, without writing a single line of code,...
- RL learning resources – normcore reads This compiles a list of links that I found useful when trying learn RL for LLMs.
- Build a link blog post Inspired by Simon Willison, I decided to put some links in my blog for helping myself to finish these links and useful resources, so that I can actually absort this.