Files
ik_llama.cpp/common
KawrakowandGitHub 7642ac3eca Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies (#2279)
* CUDA indexer topk: this is better for PP

* Don't overstep

* Cleanup

* Allow Q8_0 cache in the CUDA DSA implementation

* DS4: do not cast caches to f32

* Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies

* Re-enable -ictk | --indexer-cache-type-k
2026-08-08 17:26:59 +03:00
..
2024-07-27 07:55:01 +02:00
2026-06-04 15:43:07 +02:00
2025-12-15 08:27:20 +01:00
2026-06-14 21:07:57 -03:00
2023-11-13 14:16:23 +02:00