<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://chenboyuan.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://chenboyuan.com/" rel="alternate" type="text/html" /><updated>2026-10-07T20:34:03-07:00</updated><id>https://chenboyuan.com/feed.xml</id><title type="html">Boyuan Chen</title><subtitle>Boyuan Chen is a researcher at Huawei Canada working on LLM post-training for software engineering, distributed AI systems, and the occasional photograph.</subtitle><author><name>Boyuan Chen</name></author><entry><title type="html">从 SFT 到 RLVR 的平滑过渡：GRPO 的直觉解释</title><link href="https://chenboyuan.com/2025/12/14/sft-to-rlvr/" rel="alternate" type="text/html" title="从 SFT 到 RLVR 的平滑过渡：GRPO 的直觉解释" /><published>2025-12-14T00:00:00-08:00</published><updated>2025-12-14T00:00:00-08:00</updated><id>https://chenboyuan.com/2025/12/14/sft-to-rlvr</id><content type="html" xml:base="https://chenboyuan.com/2025/12/14/sft-to-rlvr/"><![CDATA[<p>这篇文章主要是从直觉角度来解释大模型训练中 SFT 和 RL 的关系。我们可以看到区别于SFT时“老师说的就是对”的学习方式，RL 是为了能够高效的利用“没那么正确”的样本，从而增加模型“正确”的可能性。这篇文章是我在阅读整理 Understanding Reinforcement Learning for Model Training,
and future directions with GRAPE 这篇报告的思考和笔记，也强烈建议大家去读原文。</p>

<h3 id="1-sft本质是背答案和老师越相似越好">1. SFT：本质是“背答案”，和老师越相似越好</h3>
<p>我们在 SFT（Supervised Fine-Tuning）阶段，训练的 Loss Function 本质上是<strong>负对数似然（NLL, Negative Log Likelihood）</strong>。</p>

<p>隐含的前提是：<strong>SFT 假设数据集中给出的就是“标准答案”</strong>（例如从教师模型蒸馏出的正确解题步骤和最终答案）。</p>

<p>对于第 $t$ 个 token，现有模型 $\theta$ 在当前上下文 \(\vec{s}_t\) 下生成正确 token $a_t$ 的概率是 $\pi_\theta(a_t \mid \vec{s}_t)$。我们当然希望这个概率越大越好。
对于一整句话，假设每个 token 相对独立，那么生成这句“标准答案”的概率就是所有 token 概率的乘积：（这里我们假设t从1开始，但实际上t应该从prompt后的第一个token开始）</p>

\[P(\text{Whole Sequence}) = \prod_{t=1}^{T} \pi_\theta(a_t \mid \vec{s}_t)\]

<p>我们的理论目标是让这个概率接近 1。但由于每个 token 的概率通常很小（比如 0.01），连乘之后数值会无限接近于 0，导致计算机浮点数溢出。所以我们<strong>取对数（Log）把乘法变加法，再取负号变成最小化问题</strong>，这就构成了梯度下降的 Loss：</p>

\[Loss = - \sum_{t=1}^{T} \log(\pi_\theta(a_t \mid \vec{s}_t))\]

<p>如果把一堆样本 Batch 在一起（假设有 $S$ 个样本），平均 Loss 就是：</p>

\[\mathcal{L}_{SFT} = \frac{1}{S} \sum_{i=1}^{S} NLL(\vec{text}_i, \theta)\]

<h3 id="2-reinforce给好学生加鸡腿">2. REINFORCE：给好学生“加鸡腿”</h3>
<p>SFT 的问题在于：<strong>它平等地对待了每一个样本。</strong>
但在实际生成中，有的回答是完美的，有的是一般的，甚至我们希望模型能从“负样本”中吸取教训。
于是，我们很自然地想到给 Loss 加一个<strong>权重（Reward）</strong>：</p>

\[\mathcal{L}_{RL} = \frac{1}{S} \sum_{i=1}^{S} R(\vec{text}_i) \cdot NLL(\vec{text}_i, \theta)\]

<p>这就是大名鼎鼎的 <strong>REINFORCE 算法</strong>。</p>

<p><strong>那为什么实际训练中很少直接用它？</strong>
如果 $R$ 的<strong>绝对值过大</strong>，会带来严重的<strong>梯度高方差（High Variance）</strong>。</p>
<ul>
  <li><strong>直觉解释</strong>：假设好样本得 1001 分，坏样本得 999 分。虽然分差只有 2 分，但模型在更新时，会受到一股“1001 的推力”和一股“999 的推力”。模型参数会被推得左一下、右一下，剧烈震荡，导致难以收敛。</li>
</ul>

<h3 id="3-advantage关键在于超出预期">3. Advantage：关键在于“超出预期”</h3>
<p>为了解决方差问题，我们需要对 $R$ 进行标准化处理：<strong>减去一个基线（Baseline, $V$）</strong>。
比如把 (1001, 999) 减去基线 1000，变成 (+1, -1)。这样梯度就稳定多了。</p>

<p>改进后的 Loss 公式（引入 Advantage）：</p>

\[Loss = -\frac{1}{S} \sum_{i=1}^{S} \sum_{t=1}^{T} \underbrace{\left( R(\vec{text}_i) - V_M(\vec{s}_{it}) \right)}_{\text{Advantage (A)}} \cdot \log\left( \pi_\theta(a_{it} \mid \vec{s}_{it}) \right)\]

<p>这里的 $R(\vec{text}_i)$ 是整段文本的最终得分，但基线 \(V_M(\vec{s}_{it})\) 依赖于当前时刻 $t$ 的状态（即 prompt + 已经生成的 token）。</p>

<p><strong>如何从直觉上理解 $R - V$（Advantage）？</strong>
我们可以把它看作<strong>“惊喜度”</strong>：</p>

<ol>
  <li><strong>惊喜（R &gt; V）：</strong> 假设这条样本最终得分 $R$ 很高。但在第 $t$ 个 token 之前，生成的序列平平无奇，模型觉得“这把也就拿个及格分吧”（$V$ 较低）。突然，第 $t$ 个 token 出现（比如一道数学题的关键辅助线，或者代码里的神来之笔），最终导致了高分。
    <ul>
      <li>此时 $R - V$ 很大。模型会意识到：<strong>“原来是这一步让我逆天改命的！”</strong> 于是大幅增加该 token 的概率。</li>
    </ul>
  </li>
  <li><strong>失望（R &lt; V）：</strong> 假设 $V$ 很高（模型觉得稳了），结果第 $t$ 个 token 瞎写，导致最终 $R$ 很低。
    <ul>
      <li>此时 $R - V$ 是负数。模型会反思：<strong>“原来是这一步毁了所有！”</strong> 于是抑制该 token。</li>
    </ul>
  </li>
  <li><strong>符合预期（R ≈ V）：</strong> 如果不管好坏，模型的预期 $V$ 和实际结果 $R$ 都差不多，说明没有惊喜。梯度接近 0，模型参数保持稳定，不进行无效更新。</li>
</ol>

<h3 id="4-从-ppo-到-grpo去掉了教练">4. 从 PPO 到 GRPO：去掉了“教练”</h3>
<p>搞清楚了原理，剩下的就是工程实现：<strong>$R$ 和 $V$ 怎么来？</strong></p>

<ul>
  <li><strong>获取 R（奖励）：</strong>
    <ul>
      <li><strong>RLHF：</strong> 需要训练一个 Reward Model 模仿人类偏好打分。</li>
      <li><strong>RLVR（如数学/代码）：</strong> 这个简单多了！答案对了给 1 分，格式对了给 0.2 分。规则就是 Reward。</li>
    </ul>
  </li>
  <li><strong>获取 V（基线）：</strong>
    <ul>
      <li><strong>PPO 的做法：</strong> 雇一个“教练”——<strong>训练一个独立的 Critic Model</strong>（Value Model）来实时预测当前的 $V$。这很贵，也很慢，而且 Critic 自己也需要训练。</li>
      <li><strong>GRPO 的做法：</strong> 既然我们生成了一组（Group）回答，<strong>为什么不拿“同伴的平均分”当基线？</strong>
GRPO 不需要 Critic Model。它通过采样一组输出（比如 8 个），计算这 8 个输出的平均 Reward。
        <ul>
          <li>比平均分高的，Advantage 为正（好学生）。</li>
          <li>比平均分低的，Advantage 为负（差学生）。</li>
        </ul>
      </li>
    </ul>
  </li>
</ul>

<p><strong>看到这里，聪明的你肯定会问，为什么以前我们不在一个 Batch 里直接拿所有生成结果的均值来做为 V 呢？这好像是很容易想到的呀</strong>
如果是从 PPO 过来的同学，肯定意识到，在 PPO 中，我们是把不同的 Prompt 放在一个 Batch 里，这会带来什么问题呢？不同的 Prompt 难度不一样，它们的 reward 如果放在一起，结果不会很理想。</p>

<p>假设我有一个问题的难度近似于 1+1 = 2, 而模型答对了，获得了 Reward 1.
还有一个问题是一道非常难的竞赛类题目，模型打错了，但是中间的思考过程大部分都是正确的，获得reward 0.
在这种情况下，V=0.5, 而模型朝难的竞赛题思考的过程被负向激励了（0-0.5=-0.5）久而久之，很有可能会发生的情况是，模型对难题一概不理，反正都拿到的是0的奖励。那么模型也就学不到新东西了。</p>

<p><strong>GRPO 的巧妙之处：</strong>
首先一个组内的prompt是一致的，不存在难度上的波动。它用<strong>组内平均值</strong>替代了 Critic Model 的预测值。正因如此，GRPO 计算出的 Advantage 通常是<strong>针对整条样本的</strong>（因为平均分是整条样本的平均），而不再像 PPO 那样拥有逐个 Token 的精细 $V(s)$。但对于数学推理这类“结果导向”的任务，这种简化不仅效果好，还极大地节省了显存和计算资源。</p>

<hr />
<p><em>本质上，PPO 和 GRPO 都是 REINFORCE 的变种。下一篇我们将详细介绍从 REINFORCE $\to$ TRPO $\to$ PPO $\to$ GRPO 的演进之路。Advantage 的计算当然是 GRPO 的核心，但我们还有其他的部分（如重要性采样， Clip）才能完整理解 GRPO 以及它的一系列变种如 DAPO, DR.GRPO, GSPO 等</em></p>]]></content><author><name>Boyuan Chen</name></author><category term="LLM" /><category term="RL" /><summary type="html"><![CDATA[这篇文章主要是从直觉角度来解释大模型训练中 SFT 和 RL 的关系。我们可以看到区别于SFT时“老师说的就是对”的学习方式，RL 是为了能够高效的利用“没那么正确”的样本，从而增加模型“正确”的可能性。这篇文章是我在阅读整理 Understanding Reinforcement Learning for Model Training, and future directions with GRAPE 这篇报告的思考和笔记，也强烈建议大家去读原文。]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chenboyuan.com/images/og-image.jpg" /><media:content medium="image" url="https://chenboyuan.com/images/og-image.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">A weird debugging process for Jupyter Notebook on WSL 2</title><link href="https://chenboyuan.com/2025/05/25/weird-debugging-experience/" rel="alternate" type="text/html" title="A weird debugging process for Jupyter Notebook on WSL 2" /><published>2025-05-25T00:00:00-07:00</published><updated>2025-05-25T00:00:00-07:00</updated><id>https://chenboyuan.com/2025/05/25/weird-debugging-experience</id><content type="html" xml:base="https://chenboyuan.com/2025/05/25/weird-debugging-experience/"><![CDATA[<p>事情的起因是：
在 WSL 2 下开启 Jupyter Notebook, 在 windows 下可以通过 127.0.0.1:8888 启动，但无法通过 localhost:8888 启动。</p>

<p>当时我就觉得特别奇怪，因为在我的认知里，localhost 就会被直接解析为 127.0.0.1. 于是我拿着这个问题去问 ChatGPT, Claude, 和 Gemini。最终的结论是，没有一个大模型能真正地端到端帮我解决这个问题，但是他们都或多或少给了我一些 insights 让我最终找到这个问题的根因。</p>

<p>所有的大模型，都提到了在 windows 下 ipv4 和 ipv6 的解析问题。原来，在 windows 下如果我们不去改 hosts 文件，浏览器会在localhost下优先解析 ipv6 的地址，而 jupyter notebook 没有在 ipv6 监听，于是得到 404.</p>

<p>事情真的这么简单吗？那按道理来说，只要我强制 loopback 先去解析 ipv4 地址，问题就应该解决了？或者是我索性直接让 loopback 不要去解析 ipv6，也可以？于是我先尝试了：
<code class="language-plaintext highlighter-rouge">netsh interface ipv6 set prefixpolicy ::ffff:0:0/96 60 0</code></p>

<p>这个命令是在 dual stack 系统中，让 ipv4 的地址解析优先级显著高于 ipv6，虽然我也不知道到底有没有用，但结果是当我使用 localhost:8888 时，仍然是 404. 并且我能看到在 dev tools 中，http header 的 remote addr 是<code class="language-plaintext highlighter-rouge">[::1]:8888</code>，意味着 localhost 还是被解析为 ipv6 地址，且返回了 404.</p>

<p>但是如果这个地址+端口没有任何进程在监听，chrome 根本不会显示 404 啊？它只会没有响应。</p>

<p>况且更诡异的是，在调试过程中，我打开了 Python simple httpserver 在 8000 端口，不管是 localhost 还是 127.0.0.1 都没有问题。于是我被带到这是 jupyter notebook 的问题，需要升级至 7.2.0 才能支持 dual stack 且不会拒绝远程网络的访问，等等。</p>

<p>最关键的来了，我通过 curl 命令来访问 localhost 和 127.0.0.1 （均为 ipv4）时，其实都可以访问 jupyter notebook 的服务器。</p>

<p><img src="/images/Pasted-image-20250525173534.png" alt="curl 命令关键时候很好用" /></p>

<p>这就意味着，其实只要我用 ipv4 的 localhost, jupyter 不会拒绝，没有什么 remote access 拒绝的问题。</p>

<p>在我走投无路准备放弃以后就使用 ip 地址时，我搜了搜 google，找到了这个问题：</p>

<p><a href="https://superuser.com/questions/1878272/wsl2-localhost8888-stopped-working">WSL2: localhost:8888 stopped working</a></p>

<p>看描述简直和我一模一样，于是我就换了个端口，结果就可以访问了。。。</p>

<p>于是我再接再厉，把这个 prompt 放到了 O3 里：</p>

<blockquote>
  <p>[!NOTE] 我给 O3 的最后一击
Surprisingly, if i start jupyter with another port say 9999, then both localhost and 127.0.0.1 works, why??</p>
</blockquote>

<p>O3 的回复是：</p>

<p><img src="/images/Pasted-image-20250525174008.png" alt="O3 的回复" /></p>

<p>事情到这里就很清楚了，有另一个进程一直在监听 ipv6 的 localhost … 而这个进程是——docker desktop!!!</p>

<p><a href="https://forums.docker.com/t/docker-listening-on-port-8888-after-latest-update-no-running-containers/144633/3">我也不知道 docker 为什么要设置到 8888 去</a></p>

<p>于是把 docker 的这个选相关了，就没有人再占用 8888 了，事情从此太平。</p>

<p>回头思考整个过程，我觉得大模型还是帮助了我很多。在这个过程中，我主要使用的模型是 O3，O3 现在比较 Agentic 每次会搜索很多网页来增加他回答的可信度。如果没有大模型，在这个过程中，我可能就放弃这个问题了，因为实在也没有必要。支撑我继续解决这个问题是我想看看每次我在给出反馈后大模型究竟会怎么应对。作为 Reference, 我把整个和 O3 的交流分享在<a href="https://chatgpt.com/share/68339888-5fe0-8002-8e25-6732319bae01">这里</a>.</p>]]></content><author><name>Boyuan Chen</name></author><category term="Random" /><category term="debugging" /><summary type="html"><![CDATA[事情的起因是： 在 WSL 2 下开启 Jupyter Notebook, 在 windows 下可以通过 127.0.0.1:8888 启动，但无法通过 localhost:8888 启动。]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chenboyuan.com/images/og-image.jpg" /><media:content medium="image" url="https://chenboyuan.com/images/og-image.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Thoughts on vibe coding</title><link href="https://chenboyuan.com/2025/03/24/thoughts-on-vibe-coding/" rel="alternate" type="text/html" title="Thoughts on vibe coding" /><published>2025-03-24T00:00:00-07:00</published><updated>2025-03-24T00:00:00-07:00</updated><id>https://chenboyuan.com/2025/03/24/thoughts-on-vibe-coding</id><content type="html" xml:base="https://chenboyuan.com/2025/03/24/thoughts-on-vibe-coding/"><![CDATA[<p>Recently, there’s a trend called “vibe coding,” proposed by Andrej Karpathy. It essentially means that by just telling what you want to build to LLMs, without writing a single line of code, everyone can build a standalone application. Many developers on X and Reddit are currently hyped about it, and I wanted to share my two cents.</p>

<p>Existing LLMs (e.g., Claude) and IDEs (e.g., Cursor) still have context length limitations and tend to hallucinate when problems become complex. Software engineering, on the other hand, is fundamentally about managing and controlling complexity when engineering a piece of software. This is the main reason that many “serious” developers remain skeptical about vibe coding and think the products created through vibe coding are merely toys.</p>

<p>But so what? From my own experience, even though I’m a developer with a CS degree from university, I still enjoy many benefits from vibe coding. Here are the reasons:</p>

<ol>
  <li>
    <p><strong>It drastically reduces cold-start cognitive overhead.</strong> Usually, there’s a mental struggle if I want to sit in front of my computer and start coding. Part of this is because I have a day job that’s pretty energy-consuming. And as we all know, daily energy is limited. If someone spends too much energy on their day job, it’s extremely difficult for them to start working on anything meaningful afterward. Often, they’ll end up scrolling mindlessly until bedtime. Vibe coding changes this, at least for me. When I have an idea, I no longer think, “Oh, I have to search for documentation, find out how the APIs are used, then write boilerplate classes…” Instead, vibe coding incentivizes me to START. And sometimes, simply starting can work magic.</p>
  </li>
  <li>
    <p><strong>It is an active way to learn.</strong> We often hear of the phenomenon called “tutorial hell,” where a person constantly watches YouTube videos without ever actually making anything useful. Shamefully, I used to be that kind of person. I’d watch hours of videos to “pretend” to learn some new technology, making myself feel I wasn’t wasting time. In reality, I’d forget most of the content after a few days. Now, however, I can embrace a trial-and-error process with the help of AI, and in most cases, I genuinely start picking up knowledge along the way.</p>
  </li>
  <li>
    <p><strong>It actually works.</strong> Although there are inevitably problems to solve along the way, the model doesn’t complain if you feed it error messages. Problem-solving is an essential skill in almost every field—now, we simply have an excellent helper. Previously, we had to search extensively on the internet and browse through hundreds of Stack Overflow posts (I still do sometimes).</p>
  </li>
  <li>
    <p><strong>A combination of programming and business is a killer combo.</strong> As a programmer myself, I have hobbies beyond coding. I know there are hardcore developers whose hobby is also programming, but I must admit I’m not among those prodigies. Once in a while, I also want to connect with the real world. I notice many developers being picky about apps created by non-developers, dismissing them as trivial. However, sometimes, these apps become quite popular because customers don’t care about how they’re built. In reality, many people’s needs are unmet, and programming remains a privilege limited to a few. I’m intrigued by what ideas individuals from other domains might have once they can easily create their own applications. While we often joke about the stereotypical ideas newbie indie developers choose—note-taking, weather, or finance-tracking apps—I’m confident that “outsiders” might have even more fascinating ideas than we developers could come up with.</p>
  </li>
</ol>

<p>In the end, I must acknowledge it’s still not easy. You still need to spend time sitting in front of a computer to make vibe coding work. Instead of traditional coding, you’re continuously expressing yourself through writing. As Andrej Karpathy once said, “the hottest new programming language is English.” You simply can’t bypass that effort—no one, including LLMs, can read your mind. Moving forward, I encourage everyone to pay attention to people’s real-life needs, starting by building something you’d like to have on your phone yourself. It will be fun.</p>]]></content><author><name>Boyuan Chen</name></author><category term="Essays" /><category term="random" /><summary type="html"><![CDATA[Recently, there’s a trend called “vibe coding,” proposed by Andrej Karpathy. It essentially means that by just telling what you want to build to LLMs, without writing a single line of code, everyone can build a standalone application. Many developers on X and Reddit are currently hyped about it, and I wanted to share my two cents.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chenboyuan.com/images/og-image.jpg" /><media:content medium="image" url="https://chenboyuan.com/images/og-image.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">RL learning resources - normcore reads</title><link href="https://chenboyuan.com/2025/03/18/rl-for-llm/" rel="alternate" type="text/html" title="RL learning resources - normcore reads" /><published>2025-03-18T00:00:00-07:00</published><updated>2025-03-18T00:00:00-07:00</updated><id>https://chenboyuan.com/2025/03/18/rl-for-llm</id><content type="html" xml:base="https://chenboyuan.com/2025/03/18/rl-for-llm/"><![CDATA[<p>This compiles a list of links that I found useful when trying learn RL for LLMs.</p>

<p>Links:</p>
<ul>
  <li><a href="https://github.com/wangshusen/DRL/tree/master/Notes_CN">王树森 深度强化学习</a></li>
  <li><a href="https://github.com/oreilly-japan/deep-learning-from-scratch-4/">Deep Learning from scratch 4</a></li>
  <li><a href="https://datawhalechina.github.io/easy-rl/">蘑菇书</a></li>
  <li><a href="https://github.com/alessiodm/drl-zh">drl-zh</a></li>
  <li><a href="https://github.com/huggingface/open-r1">open-r1</a></li>
  <li><a href="https://arxiv.org/pdf/2412.05265">Reinforcement Learning: An Overview 1</a></li>
</ul>]]></content><author><name>Boyuan Chen</name></author><category term="Resources" /><category term="links" /><category term="RL" /><category term="LLM" /><summary type="html"><![CDATA[This compiles a list of links that I found useful when trying learn RL for LLMs.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chenboyuan.com/images/og-image.jpg" /><media:content medium="image" url="https://chenboyuan.com/images/og-image.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Build a link blog post</title><link href="https://chenboyuan.com/2025/02/21/build-a-link-blog/" rel="alternate" type="text/html" title="Build a link blog post" /><published>2025-02-21T19:39:00-08:00</published><updated>2025-02-21T19:39:00-08:00</updated><id>https://chenboyuan.com/2025/02/21/build-a-link-blog</id><content type="html" xml:base="https://chenboyuan.com/2025/02/21/build-a-link-blog/"><![CDATA[<p>Inspired by Simon Willison, I decided to put some links in my blog for helping myself to finish these links and useful resources, so that I can actually absort this.</p>

<p>Recently, my main focus is on training LLMs for coding/software engineering. Hence the following links might be very helpful:</p>

<ol>
  <li><a href="https://rlhfbook.com/">https://rlhfbook.com/</a></li>
  <li><a href="https://github.com/therealoliver/Deepdive-llama3-from-scratch/">https://github.com/therealoliver/Deepdive-llama3-from-scratch/</a></li>
  <li><a href="https://huggingface.co/spaces/nanotron/ultrascale-playbook/">https://huggingface.co/spaces/nanotron/ultrascale-playbook/</a></li>
  <li><a href="https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1/">https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1/</a></li>
  <li><a href="https://github.com/Open-Reasoner-Zero/Open-Reasoner-Zero/blob/main/ORZ_paper.pdf/">https://github.com/Open-Reasoner-Zero/Open-Reasoner-Zero/blob/main/ORZ_paper.pdf/</a></li>
  <li><a href="https://www.bilibili.com/video/BV1sd4y167NS/?spm_id_from=333.337.search-card.all.click&amp;vd_source=30dd777f9dac71e1bdc7362c649969f6/">https://www.bilibili.com/video/BV1sd4y167NS/?spm_id_from=333.337.search-card.all.click&amp;vd_source=30dd777f9dac71e1bdc7362c649969f6/</a></li>
  <li><a href="https://docs.google.com/presentation/d/11KWCKUORnPpVMSY6vXgBeFSWo7fJcuGQ9yuR6vC1pzE/edit#slide=id.p/">https://docs.google.com/presentation/d/11KWCKUORnPpVMSY6vXgBeFSWo7fJcuGQ9yuR6vC1pzE/edit#slide=id.p/</a></li>
</ol>

<p>These are mostly related to RL and Infra.</p>]]></content><author><name>Boyuan Chen</name></author><category term="Resources" /><category term="links" /><summary type="html"><![CDATA[Inspired by Simon Willison, I decided to put some links in my blog for helping myself to finish these links and useful resources, so that I can actually absort this.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://chenboyuan.com/images/og-image.jpg" /><media:content medium="image" url="https://chenboyuan.com/images/og-image.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>