Default Branch

592feef04a · talk-llama : sync llama.cpp · Updated 2026-08-07 22:59:49 +04:00

Branches

267e15a46d · cuda : avoid async allocs in CUDA mel code · Updated 2024-06-12 10:52:15 +04:00    LLM-Inference

3583
1

5801b8ac64 · cuda : fix HIPBLAS build · Updated 2024-06-11 20:13:43 +04:00    LLM-Inference

3584
1

13c5446759 · Update ggml-cuda/mmvq.cu · Updated 2024-06-11 18:37:32 +04:00    LLM-Inference

3586
2

059bcd3009 · ci : fix CUDA builds · Updated 2024-06-11 12:40:19 +04:00    LLM-Inference

3586
1

ba69578828 · whisper : add whisper_token_count helper · Updated 2024-03-25 16:46:07 +04:00    LLM-Inference

3730
2

66df44b0b7 · alloc : fix allocation data of pre-allocated leafs · Updated 2024-03-16 18:47:14 +04:00    LLM-Inference

3739
2

f25edade2b · whisper : alternative way to handle the external encoders · Updated 2024-02-12 18:32:26 +04:00    LLM-Inference

3866
2

15c4fdce45 · chess : tuning performance · Updated 2023-11-30 12:50:47 +04:00    LLM-Inference

4096
21

4260d4fc70 · wchess : minor · Updated 2023-11-28 17:10:18 +04:00    LLM-Inference

4096
11

c8b3bc6a0d · cuda : use CUBLAS_COMPTE_F32 insted of CUBLAS_COMPUTE_F16 · Updated 2023-11-27 13:57:07 +04:00    LLM-Inference

4078
1

ee2971bf6a · bench : multi-thread memcpy · Updated 2023-11-21 23:57:07 +04:00    LLM-Inference

4096
1

ec96d68402 · whisper : quantize encoder only · Updated 2023-11-16 18:19:02 +04:00    LLM-Inference

4106
1

270b1e48db · cuda : sync llama.cpp fixes · Updated 2023-11-15 17:52:06 +04:00    LLM-Inference

4118
14

5031f54717 · whisper : try to fix the parallel whisper_state functionality (#1479) · Updated 2023-11-12 16:52:38 +04:00    LLM-Inference

4126
21

a2f3b82db3 · whisper : free backend instances in whisper_state · Updated 2023-11-12 16:31:51 +04:00    LLM-Inference

4126
23

7a91a3ba60 · bench-all : add q4 models · Updated 2023-11-11 00:23:18 +04:00    LLM-Inference

4126
16

bf4110dbcf · whisper : wip sched (not working yet) · Updated 2023-11-09 21:07:54 +04:00    LLM-Inference

4131
2

40be74271f · models : update readme · Updated 2023-11-07 15:53:01 +04:00    LLM-Inference

4137
4

aaa3b5e5f6 · ggml : try to fix the abort mechanism · Updated 2023-11-05 22:02:24 +04:00    LLM-Inference

4144
1

673c55c683 · whisper : print log when using distilled models · Updated 2023-11-05 21:43:04 +04:00    LLM-Inference

4147
2