Default Branch

84e908c625 · ci: fix thread sanitizer + remove ccache (#26927) · Updated 2026-08-12 20:01:12 +04:00

Branches

865066621b · llama.swiftui : improve bench · Updated 2023-12-17 21:37:22 +04:00    LLM-Inference

8756
12

f86b9d152c · lookup : minor · Updated 2023-12-17 19:25:28 +04:00    LLM-Inference

8754
9

d2f1e0dacc · Merge branch 'cuda-cublas-opts' into gg/phi-2 · Updated 2023-12-17 10:41:46 +04:00    LLM-Inference

8752
17

b0547d2196 · gguf-py : fail fast on nonsensical special token IDs · Updated 2023-12-16 03:06:42 +04:00    LLM-Inference

8754
1

c8554b80be · Merge branch 'master' of https://github.com/ggerganov/llama.cpp into ceb/fix-cuda-warning-flags · Updated 2023-12-13 21:06:01 +04:00    LLM-Inference

8766
12

e1241d9b46 · metal : switch to execution barriers + fix one of the barriers · Updated 2023-12-13 15:56:45 +04:00    LLM-Inference

8777
47

fc5f334689 · readme : add API change notice · Updated 2023-12-07 14:35:02 +04:00    LLM-Inference

8779
15

af99c6fbfc · llama : remove memory_f16 and kv_f16 flags · Updated 2023-12-05 20:18:16 +04:00    LLM-Inference

8791
26

3cb1c348b3 · metal : try to improve batched decoding · Updated 2023-12-02 00:01:58 +04:00    LLM-Inference

8796
2

eb594c0f7d · alloc : fix build with debug · Updated 2023-12-01 12:46:05 +04:00    LLM-Inference

8820
14

5b74310e6e · build : enable libstdc++ assertions for debug builds · Updated 2023-12-01 03:18:24 +04:00    LLM-Inference

8805
1

bb39b87964 · ggml : restore abort() in GGML_ASSERT · Updated 2023-11-28 04:27:09 +04:00    LLM-Inference

8824
1

87f4102a70 · llama : revert n_threads_batch logic · Updated 2023-11-27 23:47:35 +04:00    LLM-Inference

8825
3

6272b6764a · use stride=128 if built for tensor cores · Updated 2023-11-27 22:09:14 +04:00    LLM-Inference

8828
3

8d8b76d469 · lookahead : add comments · Updated 2023-11-26 13:26:55 +04:00    LLM-Inference

8840
9

21b70babf7 · straightforward /v1/models endpoint · Updated 2023-11-24 20:22:39 +04:00    LLM-Inference

8841
12

f8e9f11428 · common : add -dkvc arg for enabling kv cache dumps · Updated 2023-11-23 20:47:56 +04:00    LLM-Inference

8847
4

f824902623 · YaRN : correction to GPT-NeoX implementation · Updated 2023-11-16 02:10:52 +04:00    LLM-Inference

8879
1

d0445a2eff · better documentation · Updated 2023-11-10 04:38:20 +04:00    LLM-Inference

8896
3

47d604fa2d · fix issues · Updated 2023-11-05 16:20:22 +04:00    LLM-Inference

8910
3