{"author":"pythongiant","children":[{"author":"pythongiant","children":[],"created_at":"2026-07-07T11:54:38.000Z","created_at_i":1783425278,"id":48816465,"options":[],"parent_id":48816446,"points":null,"story_id":48816446,"text":"Hey guys, i&#x27;m especially interested in feedback on the kernel design and integration onto mlx lm.<p>Its on pypi as well as a simple pip install mlx-turboquant :P","title":null,"type":"comment","url":null}],"created_at":"2026-07-07T11:52:04.000Z","created_at_i":1783425124,"id":48816446,"options":[],"parent_id":null,"points":1,"story_id":48816446,"text":"Hi HN,<p>I built mlx-turboquant, an implementation of Google&#x27;s TurboQuant KV-cache compression algorithm for Apple&#x27;s MLX framework.<p>The repository includes quality benchmarks, memory benchmarks, and a modular implementation so individual pieces (PolarQuant, QJL, packing, codebooks) can be studied independently.","title":"Show HN: TurboQuant for mlx-lm (Apple Silicon)","type":"story","url":"https://github.com/pythongiant/mlx_turboquant"}
