Paper Digest: 2026-09-10
今天最值得看的 paper,我会选 Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails。 论文研究一个很容易被低估的问题:当 agent harness 已经针对某个模型优化过,再用强模型的成功轨迹微调这个弱模型,效果会怎样? 答...
今天最值得看的 paper,我会选 Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails。 论文研究一个很容易被低估的问题:当 agent harness 已经针对某个模型优化过,再用强模型的成功轨迹微调这个弱模型,效果会怎样? 答...
今天最值得看的 paper,我会选 Miles v0.1: Production-Level Post-Training。 Miles 是一套面向 frontier post-training 的开源 full-stack system。它把 rollout、trainer、weight synchronization 和异步执行放进同一套架构,并用一个足够有分量的 case study ...
今天最值得看的 paper,我会选 Unlocking Lossless Speedups in LLMs via Discrete Diffusion。 作者提出了一类 diffusion-augmented LLM,并将模型命名为 Uno。它保留标准 autoregressive LLM 的训练和输出分布,同时增加一组轻量 diffusion weights,一次并行提出多个 toke...
今天最值得看的 paper,我会选 Iris: Climbing to the Search Frontier。 Iris 训练了两个 search agent:35B-A3B 的 Iris-mini,以及 397B-A17B 的 Iris-pro。论文的价值不只在 benchmark 分数。它把 task synthesis、trajectory filtering、SFT、online...
今天最值得看的 paper,我会选 Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments。 训练 code agent 时,trajectory 已经积累了很多,真正稀缺的是可以反复执行、产生反馈、继续提出新任务的 environment。一条 trajectory 只记录一次 i...
今天最值得看的 paper,我会选 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills。 Agent 做 ML research 时经常卡在一些很具体的地方:某个 package 的 API 怎么组合,training config 有哪些隐含约束,evaluation script 如何跑通,失败后该检查什么。模...
今天最值得看的 paper,我会选 From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix。 大公司的内部 LLM 平台很容易长出一个 model zoo。两百多个应用各自选择当时最合适的模型,新模型持续上线,旧模型又因为迁移成本迟迟...
今天最值得看的 paper,我会选 Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement。 On-policy distillation(OPD)看起来像一种更密集的 RLVR。Student 先采样 trajectory,teacher 再给每个 token 提供 advant...
今天最值得看的 paper,我会选 LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering。 长时间运行的 coding agent,真正困难的部分经常发生在每一轮 coding 之间。 当前实现做到什么程度?进度说明有没有过期?某个局部 test 通过,是否足以证明整个任务完成?下一轮应该继续...
今天最值得看的 paper,我会选 TTPO: Test-Time Policy Optimization。 它研究一个很棘手的问题:没有 ground-truth label,模型还能不能在 test time 继续训练自己? 最直接的办法是对同一道题采样很多条 rollout,用多数票当 pseudo-label。可是在 competition math 这种高难度数据上,多数票经常...