mirror of
https://github.com/ikawrakow/ik_llama.cpp.git
synced 2026-08-12 22:29:39 +04:00
* CUDA indexer topk: this is better for PP * Don't overstep * Cleanup * Allow Q8_0 cache in the CUDA DSA implementation * DS4: do not cast caches to f32 * Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies * Re-enable -ictk | --indexer-cache-type-k