Default Branch

26ceed9d40 · CUDA: clear MMQ row padding on partially offloaded quantized weights (#2292) · Updated 2026-08-11 11:07:01 +04:00

Branches

140c72c735 · Muse-glimmer: Slightly better split mode graph (+2% TG) · Updated 2026-08-12 12:21:32 +04:00

2
8

382d529d70 · DSA: do not copy V rows when V == K · Updated 2026-08-10 20:47:02 +04:00

4
1

cf70630b37 · Actually fix quantized indexer cache on CUDA · Updated 2026-08-10 17:04:59 +04:00

6
1

22f28e819a · Indexer topk: on the CPU repack Q8_0 indexer cache · Updated 2026-08-10 13:34:56 +04:00

9
1

2610fb63da · Allow Q8_0 cache in the CUDA DSA implementation · Updated 2026-08-08 17:54:29 +04:00

15
4

b8d24c9c42 · DS4: do not cast caches to f32 · Updated 2026-08-08 17:54:29 +04:00

15
5

f7f0982c7b · Re-enable -ictk | --indexer-cache-type-k · Updated 2026-08-08 17:54:29 +04:00

15
7

57ceaf0e72 · Cleanup · Updated 2026-08-08 17:54:28 +04:00

15
3

d7836010a1 · Reduce the indexer temporary buffer size · Updated 2026-08-07 18:05:30 +04:00

19
1

fc309e83a0 · Minor · Updated 2026-08-07 14:39:11 +04:00

19
2

18fe89a003 · Do not include ggml-impl.h in ggml-cuda.cu · Updated 2026-08-06 15:45:44 +04:00

28
1

fb5fb74880 · Better placement of MoE tensors with -ncmoe and 1 GPU · Updated 2026-08-06 10:52:57 +04:00

28
1

5c8db78a84 · Fix CUDA silu kernel for merged up/gate with limit · Updated 2026-08-06 09:18:57 +04:00

29
4

7bfd4cd708 · Cleanup · Updated 2026-08-05 09:03:44 +04:00

30
2

2015e6262c · Let's tell the user what we did · Updated 2026-08-04 20:24:07 +04:00

35
2

0456123b34 · Fix #2201 · Updated 2026-08-04 18:56:46 +04:00

35
1

862828d02a · Revert "DS4: faster long-context TG (#2201)" · Updated 2026-08-04 11:02:51 +04:00

39
1

79176e3b32 · Do not quantize integer tensors · Updated 2026-08-03 11:36:27 +04:00

40
1

6b1b43fe9d · Allow K to be f32 in ggml_cuda_op_indexer_topk · Updated 2026-08-02 20:24:04 +04:00

46
3

ae08cb1cea · Disable quantized indexer cache · Updated 2026-08-02 18:59:26 +04:00

45
1