Default Branch

592feef04a · talk-llama : sync llama.cpp · Updated 2026-08-07 22:59:49 +04:00

Branches

3ac0558009 · ios : update SPM package · Updated 2023-09-15 13:13:33 +04:00    LLM-Inference

4176
44

09a6325de5 · ggml : use sched_yield when using BLAS + add comment · Updated 2023-09-12 14:33:09 +04:00    LLM-Inference

4177
2

3c50be2217 · whisper : remove comment · Updated 2023-09-10 14:27:06 +04:00    LLM-Inference

4194
9

9c1a414feb · use secrets.GPG_PRIVATE_KEY and GPG_PASSPHRASE · Updated 2023-09-09 18:49:13 +04:00    LLM-Inference

4179
5

3efe146d27 · ci : enable java package publishing · Updated 2023-08-30 23:13:38 +04:00    LLM-Inference

4201
1

8cbc363561 · coreml : attempt to fix ANE-optimized models · Updated 2023-07-12 00:03:53 +04:00    LLM-Inference

4239
1

c6174cb868 · wip · Updated 2023-05-02 22:47:12 +04:00    LLM-Inference

4307
1

c456ca476b · llama podcast · Updated 2023-04-01 14:13:27 +04:00    LLM-Inference

4375
1

3627ef51f6 · minor · Updated 2023-03-27 00:54:52 +04:00    LLM-Inference

4392
7

0244810697 · rebase on master after whisper_state changes · Updated 2023-03-26 17:09:06 +04:00    LLM-Inference

4392
3

4f074fb7a8 · tmp : demonstrate how to measure time of ggml ops · Updated 2023-03-09 11:28:06 +04:00    LLM-Inference

4403
1

a0da7f71a2 · command : wip in progress, improve guided decoding · Updated 2023-02-19 21:39:05 +04:00    LLM-Inference

4419
1

ec44ad0a75 · diarization : try conv and self-attention embeddings · Updated 2023-02-19 15:00:12 +04:00    LLM-Inference

4420
4

59c997ca2d · wip ignore · Updated 2023-02-15 21:11:12 +04:00    LLM-Inference

4427
1

7aa1174315 · bench : fix Windows linkage by moving ggml benches in whisper lib .. · Updated 2023-01-18 23:16:25 +04:00    LLM-Inference

4465
1

e2aa556a99 · whisper : experiments with Flash Attention in the decoder · Updated 2023-01-07 23:00:51 +04:00    LLM-Inference

4487
1

4e6d2e98ab · ggml : try to improve threading · Updated 2022-12-29 15:05:20 +04:00    LLM-Inference

4527
1

683f111088 · ggml : initial tests with libnvblas · Updated 2022-12-09 00:01:52 +04:00    LLM-Inference

4606
1

e0bd97f41f · ggml : use macros to inline FP16 <-> FP32 conversions · Updated 2022-12-07 00:05:33 +04:00    LLM-Inference

4610
1

0a2621b637 · stream : add "max_tokens" cli arg · Updated 2022-11-20 23:22:02 +04:00    LLM-Inference

4688
5