Default Branch

84e908c625 · ci: fix thread sanitizer + remove ccache (#26927) · Updated 2026-08-12 20:01:12 +04:00

Branches

bb9abb5cd8 · imatrix: guard Q4_0/Q5_0 against ffn_down craziness · Updated 2024-01-16 11:56:05 +04:00    LLM-Inference

8518
2

9998ecd191 · llama : add phixtral support (wip) · Updated 2024-01-13 16:24:07 +04:00    LLM-Inference

8548
1

1fb563ebdc · py : try to fix flake stuff · Updated 2024-01-13 15:42:35 +04:00    LLM-Inference

8549
2

9bfcb16fd3 · Add llama enum for IQ2_XS · Updated 2024-01-11 20:24:12 +04:00    LLM-Inference

8598
11

24096933b0 · server : try to fix infill when prompt is empty · Updated 2024-01-09 13:27:29 +04:00    LLM-Inference

8600
1

7216af5c09 · ggml : fix 32-bit ARM compat (cont) · Updated 2024-01-09 12:33:16 +04:00    LLM-Inference

8603
2

d57cb9c294 · passkey : add readme · Updated 2024-01-08 13:13:44 +04:00    LLM-Inference

8613
7

7cfde78190 · llama : remove redundant GQA check · Updated 2024-01-06 18:04:20 +04:00    LLM-Inference

8621
1

9f51f3e695 · metal : opt mul_mm_id · Updated 2024-01-02 22:50:18 +04:00    LLM-Inference

8647
17

4cc78d3873 · ggml : force F32 precision for ggml_mul_mat · Updated 2024-01-02 19:54:56 +04:00    LLM-Inference

8646
1

b5af7ad84f · llama : refactor quantization to avoid <mutex> header · Updated 2024-01-02 17:56:57 +04:00    LLM-Inference

8649
1

120a1a5515 · llama : auto download HF models if URL provided · Updated 2024-01-02 15:29:06 +04:00    LLM-Inference

8650
1

f64e4f04e7 · ggml : testing GPU FP precision via quantized CPY · Updated 2023-12-30 21:11:40 +04:00    LLM-Inference

8668
1

f32f30bc57 · test · Updated 2023-12-26 19:52:42 +04:00    LLM-Inference

8698
1

ab1b75166f · Merge branch 'master' into gg/ggml_scale · Updated 2023-12-22 00:35:11 +04:00    LLM-Inference

8721
4

7c87353e61 · common : remove incorrect --model-draft default · Updated 2023-12-21 21:17:12 +04:00    LLM-Inference

8729
1

a40f6110f0 · ggml : force F32 precision for ggml_mul_mat · Updated 2023-12-19 18:34:59 +04:00    LLM-Inference

8736
1

3c734f4941 · plamo : testing · Updated 2023-12-18 19:06:05 +04:00    LLM-Inference

8741
13

a462159c43 · cuda : ggml_cuda_op_mul_mat_cublas support F32 precision · Updated 2023-12-18 16:24:29 +04:00    LLM-Inference

8741
16

1b05817112 · decode : fix logits_valid for old API · Updated 2023-12-18 03:49:21 +04:00    LLM-Inference

8742
1