Default Branch

84e908c625 · ci: fix thread sanitizer + remove ccache (#26927) · Updated 2026-08-12 20:01:12 +04:00

Branches

3ef358fffd · Revert "cuda : use CUDA memory pool with async memory allocation/deallocation when available (#3903)" · Updated 2023-11-05 00:26:51 +04:00    LLM-Inference

8914
2

46868a499e · metal : multi-simd softmax · Updated 2023-11-01 23:16:34 +04:00    LLM-Inference

8939
1

a8796f9609 · llm : cleanup + comments · Updated 2023-11-01 22:08:02 +04:00    LLM-Inference

8948
4

7420bef83e · wip wip wip · Updated 2023-11-01 10:51:43 +04:00    LLM-Inference

8948
1

afb3929279 · Merge branch 'master' into llama-refactor · Updated 2023-10-31 22:35:31 +04:00    LLM-Inference

8950
21

29fe516913 · wip · Updated 2023-10-31 20:36:37 +04:00    LLM-Inference

8951
1

dab42893c9 · scripts : working curl pipe · Updated 2023-10-31 19:03:56 +04:00    LLM-Inference

8951
3

7923b70cb8 · llama : add llm_build_inp_embd helper · Updated 2023-10-31 18:43:08 +04:00    LLM-Inference

8956
37

4b3cb98d46 · ggml-impl : move extern "C" to start of file · Updated 2023-10-30 21:05:58 +04:00    LLM-Inference

8952
7
lto

bc28aaa8c2 · make : use -lfto=auto to avoid warnings and maintain perf · Updated 2023-10-30 18:00:53 +04:00    LLM-Inference

8952
5

15267192c0 · llama : refactor tensor offloading as callback · Updated 2023-10-29 15:04:36 +04:00    LLM-Inference

8956
15

8a86b95e87 · quantize : --pure option for disabling k-quant mixtures · Updated 2023-10-29 00:37:03 +04:00    LLM-Inference

8957
3

de7e0912b6 · convert : ignore tokens if their IDs are within [0, vocab_size) · Updated 2023-10-28 16:01:36 +04:00    LLM-Inference

8960
1

bbfc62ac2f · sampling : temp == 0.0 -> no probs, temp < 0.0 -> probs · Updated 2023-10-28 15:04:57 +04:00    LLM-Inference

8968
3

cd3e20fb50 · cuda : fix multi-gpu with tensor cores · Updated 2023-10-28 00:11:50 +04:00    LLM-Inference

8967
3

49af767fad · build : add compile option to force use of MMQ kernels · Updated 2023-10-27 14:21:04 +04:00    LLM-Inference

8969
7

d798a17c34 · cuda : add TODO for calling cublas from kernel + using mem pool · Updated 2023-10-24 17:33:24 +04:00    LLM-Inference

8983
10

6966474928 · cuda : play with faster Q4_0 dequantization · Updated 2023-10-24 11:29:40 +04:00    LLM-Inference

8983
8

b9bb4cbe86 · Separate bug and enhancement template + no default title · Updated 2023-10-23 19:59:11 +04:00    LLM-Inference

8983
1

c0f4d54870 · server : add comment about changing slot_state to bool · Updated 2023-10-22 23:24:39 +04:00    LLM-Inference

8989
72