谈谈LLM生成文本的惩罚参数
之前在 一文搞懂大模型生成文本的解码策略 中简单介绍过惩罚参数(重复惩罚、频率惩罚、存在惩罚),本文来详谈一下这三者的区别以及代码实现。重复惩罚(repetition_penalty)工作机制: 直接针对当前上下文中(包括输入+已生成的token)已经出现过的 tok…
English triage
Chinese technical note: 谈谈LLM生成文本的惩罚参数
吃果冻不吐果冻皮 published a Zhihu article relevant to inference and AI systems, frontier and open model development. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.