Default Branch

84e908c625 · ci: fix thread sanitizer + remove ccache (#26927) · Updated 2026-08-12 20:01:12 +04:00

Branches

91c453fb11 · One cannot possibly be defining static_assert in a C++ compilation · Updated 2024-02-05 15:22:14 +04:00    LLM-Inference

8331
2

49a483e0f2 · wip · Updated 2024-02-04 14:34:36 +04:00    LLM-Inference

8357
60

a647257b47 · cuda : express strides with helper constants · Updated 2024-02-04 13:45:26 +04:00    LLM-Inference

8357
60

b957b8f5f6 · cuda : add flash_attn kernel (wip) · Updated 2024-02-01 21:49:57 +04:00    LLM-Inference

8361
39

ac26f27028 · cuda : increase C to 128 for better performance · Updated 2024-02-01 19:08:29 +04:00    LLM-Inference

8361
61

1ad42b1f1e · ggml : ggml_soft_max uses F16 mask · Updated 2024-01-31 22:33:59 +04:00    LLM-Inference

8361
36

719a087138 · iq3_xxs: forgotten update of the grid points · Updated 2024-01-30 20:39:07 +04:00    LLM-Inference

8375
1

2bf91c5306 · metal : clean up · Updated 2024-01-25 15:29:45 +04:00    LLM-Inference

8471
23

6ccbd1777a · wip · Updated 2024-01-24 17:45:04 +04:00    LLM-Inference

8471
18

da23b56f25 · wip : no ic 8 step · Updated 2024-01-24 15:25:34 +04:00    LLM-Inference

8471
18

06c2d0d117 · wip · Updated 2024-01-24 00:42:43 +04:00    LLM-Inference

8471
14

a9681febd6 · ggml : online attention (CPU) · Updated 2024-01-20 18:45:41 +04:00    LLM-Inference

8471
4

32a392fe68 · try a differerent fix · Updated 2024-01-20 02:10:23 +04:00    LLM-Inference

8472
2

4a3bc1522e · py : linting with mypy and isort · Updated 2024-01-20 00:18:58 +04:00    LLM-Inference

8473
3

1453215165 · kompute : fix ggml_add kernel · Updated 2024-01-19 02:09:16 +04:00    LLM-Inference

8589
105

ccc78a200e · hellaswag: speed up even more by parallelizing log-prob evaluation · Updated 2024-01-18 20:25:29 +04:00    LLM-Inference

8489
1

2917e6b528 · Merge branch 'master' into gg/imatrix-gpu-4931 · Updated 2024-01-17 20:43:45 +04:00    LLM-Inference

8496
10

23742deb5b · py : fix padded dummy tokens (I hope) · Updated 2024-01-17 17:44:22 +04:00    LLM-Inference

8515
4

9fd1e83f6d · Use Q4_K for attn_v for Q2_K_S when n_gqa >= 4 · Updated 2024-01-17 14:16:08 +04:00    LLM-Inference

8501
1

49bafe0986 · tests : avoid creating RNGs for each tensor · Updated 2024-01-17 12:40:55 +04:00    LLM-Inference

8504
6