Hi!
Have you run any dedicated throughput benchmarks for prefill and decode speeds using the hq-e8-2b KV codec? The memory density is phenomenal (~7x over bf16), but I'm curious if the E8-lattice decoding and Rice coding overhead noticeably impacts the attention kernels (tok/s) compared to standard nvfp4 or rk2v4-e8.
Hi!
Have you run any dedicated throughput benchmarks for prefill and decode speeds using the hq-e8-2b KV codec? The memory density is phenomenal (~7x over bf16), but I'm curious if the E8-lattice decoding and Rice coding overhead noticeably impacts the attention kernels (tok/s) compared to standard nvfp4 or rk2v4-e8.