Google Research·· 2026-03-25精选AI 评分74
Google Research 发布 TurboQuant,以极端压缩降低 AI 内存开销
TurboQuant: Redefining AI efficiency with extreme compression
AI 导读
Google Research 介绍 TurboQuant,以及配套的 QJL 和 PolarQuant 压缩算法,用于降低 LLM 的 KV cache 和高维向量搜索的内存开销。
推荐理由
文章系统介绍了 TurboQuant、QJL 和 PolarQuant 的压缩机制与实验结果,展示其在 KV cache 和向量搜索中的内存与速度收益。
来源:Google Research · research.google