Default Branch

84e908c625 · ci: fix thread sanitizer + remove ccache (#26927) · Updated 2026-08-12 20:01:12 +04:00

Branches

cb79f8a2d8 · llama : add SKIP_KQ_KQV option · Updated 2023-10-22 10:58:29 +04:00    LLM-Inference

8989
3

56ba00b923 · sampling : hide prev behind API and apply #3661 · Updated 2023-10-20 19:53:27 +04:00    LLM-Inference

8992
6

ad2727d091 · Merge branch 'master' into speculative-tree · Updated 2023-10-18 11:50:58 +04:00    LLM-Inference

9003
18

932589c0ef · Honor -ngl option for Cuda offloading in llava · Updated 2023-10-14 04:12:10 +04:00    LLM-Inference

9017
1

5261aee8d8 · sampling : one sequence per sampling context · Updated 2023-10-12 21:36:44 +04:00    LLM-Inference

9020
1

2fcdf869cd · batched-bench : add mmq CLI arg · Updated 2023-10-11 20:42:33 +04:00    LLM-Inference

9032
7

ee7456926e · ggml-alloc : fix assert in debug builds · Updated 2023-10-09 16:33:12 +04:00    LLM-Inference

9041
1

ee268b5446 · llama : no longer perform uninitialized access to the KV cache · Updated 2023-10-08 12:49:38 +04:00    LLM-Inference

9048
5

acead654d2 · Merge branch 'master' into fix-refact · Updated 2023-10-08 12:25:16 +04:00    LLM-Inference

9048
4

6b9554a740 · metal : print more GPU info + disable mul_mm for MTLGPUFamiliy < Apple7 · Updated 2023-10-08 10:55:13 +04:00    LLM-Inference

9055
5

ba44776dc2 · bump version · Updated 2023-10-07 22:47:48 +04:00    LLM-Inference

9054
6

5ab6c2132a · server-parallel : add "--reverse-prompt" + compiler warning fixes · Updated 2023-10-06 15:32:19 +04:00    LLM-Inference

9067
4

5418932b71 · llama : fix comments for llama_kv_cache API · Updated 2023-10-03 22:01:52 +04:00    LLM-Inference

9092
5

c5650ed470 · server : avoid context swaps by shifting the KV cache · Updated 2023-09-28 20:03:36 +04:00    LLM-Inference

9116
57

72e7ef4e53 · simple : fixes · Updated 2023-09-27 01:19:36 +04:00    LLM-Inference

9142
48

784d14ed31 · llama : store non-RoPEd K cache (WIP) · Updated 2023-09-18 00:43:07 +04:00    LLM-Inference

9154
5

92a4f86879 · llama : make starcoder graph build more consistent with others · Updated 2023-09-15 18:57:10 +04:00    LLM-Inference

9164
20

e7e7b11455 · llama : remove experimental stuff · Updated 2023-09-14 23:52:01 +04:00    LLM-Inference

9176
3

2f689dee06 · metal : minor · Updated 2023-09-07 16:33:21 +04:00    LLM-Inference

9209
5

30ac7a4117 · gitignore : metal · Updated 2023-09-04 23:23:16 +04:00    LLM-Inference

9221
12