Default Branch

c46ffaa566 · fix dspark: seed draft block at id_last's true position (+1 off-by-one) (#2296) · Updated 2026-08-12 20:21:30 +04:00

Branches

6b1b43fe9d · Allow K to be f32 in ggml_cuda_op_indexer_topk · Updated 2026-08-02 20:24:04 +04:00

46
3

ae08cb1cea · Disable quantized indexer cache · Updated 2026-08-02 18:59:26 +04:00

45
1

a2f85b65f9 · Fix IQ4_NL_R4 GEMM on CPUs with FANCY_SIMD enabled · Updated 2026-08-02 11:48:35 +04:00

48
1

b2c8508a59 · Bucket top_k (CPU): ~3% better TG at 128k context · Updated 2026-08-01 10:29:01 +04:00

51
1

aab3e1579b · Fix IQ3_XXS CPU GEMM · Updated 2026-08-01 09:40:00 +04:00

56
1

109ac60f1c · Fix constants.py · Updated 2026-07-31 13:58:21 +04:00

57
1

e8649de622 · Don't make a copy when you don't need to · Updated 2026-07-31 08:19:51 +04:00

59
2

d2c040d15b · Update README and CONTRIBUTING · Updated 2026-07-30 20:23:53 +04:00

59
1

a035b315de · Set n_gpu_layers automatically · Updated 2026-07-30 19:29:28 +04:00

62
1

d4dfbf44a9 · Minor · Updated 2026-07-30 17:59:51 +04:00

66
2

48292aefab · Update links · Updated 2026-07-30 16:52:51 +04:00

65
1

c0c0865be8 · Option to turn it off at compile time · Updated 2026-07-30 14:14:52 +04:00

67
2

2c1890bcb5 · Also this · Updated 2026-07-29 16:08:31 +04:00

68
2

393a362d96 · Revert CUDA concat change in #2179 · Updated 2026-07-28 13:17:20 +04:00

70
1

23eced923c · Remove commented out code · Updated 2026-07-28 10:18:30 +04:00

70
3

9799fc0377 · Add AVX512 implementation for MXFP4_R8 · Updated 2026-07-27 17:52:46 +04:00

72
3

18089273b4 · DS4 refactoring (cont'd) · Updated 2026-07-27 13:05:54 +04:00

72
1

d2074fad25 · Minor · Updated 2026-07-27 10:11:45 +04:00

75
2

18ac719c27 · Increase max. number of graph splitinputs to 64 · Updated 2026-07-25 20:27:03 +04:00

80
4

922ba730a0 · Fix quantized cache · Updated 2026-07-24 17:49:57 +04:00

82
8