在SFT训练阶段,有大量短文本数据与少量的长文本数据,从训练效率和模型性能的角度,应该如何选择批处理策略?
在SFT训练阶段,有大量短文本数据与少量的长文本数据,从训练效率和模型性能的角度,应该如何选择批处理策略?1、sorted batching 在传统的批处理(batching)中,由于序列长度不一,通常需要通过填充(padding)将同一批次内的所有序列对齐到最大长度,这…
English triage
Chinese technical note: 在SFTtraining阶段,有大量短文本数据与少量的长文本数据,从training效率和模型性能的角度,应该如何选择批处理策略?
xhchen published a Zhihu article relevant to post-training and RL. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.