为什么加速LLM推断有KV Cache而没有Q Cache?
English triage
Chinese technical note: Why加速LLM推断有KV Cache而没有Q Cache?
小冬瓜AIGC published a Zhihu answer relevant to frontier and open model development. The original Chinese excerpt is included below so the feed can preserve the raw source while giving English readers enough context to triage the item.