mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-12 22:31:11 +04:00
f9e832c10e9444cb168ddcb579cc62c154f3068b
* server: don't walk Windows junctions in file_glob_search std::filesystem reports a junction as a plain directory, so the symlink guard misses it and a junction pointing back at an ancestor is walked until the path length gives out read the reparse tag and treat a symlink and a mount point as links, leaving any other reparse point walkable so cloud placeholders and dedup stubs still get searched look junk directory names up case insensitively on Windows, where NTFS makes Build the same directory as build test that a junk directory stays selectable while its contents stay out of search results * server: report a directory the walk could not read a directory that fails to open or to iterate was skipped in silence, so a caller got a listing that looked complete while a whole subtree was missing: a path over the platform limit, a volume going away, a name the filesystem rejects skip_permission_denied never reaches this path, so an error here is an incomplete answer rather than a deliberate omission, and it now sets the truncated flag * server: simplify the file_glob_search listing plumbing return a small result struct instead of two out params and a caller path that only fed an error string, taking list_entries from six parameters down to three scope the error code to the directory being read, act on the status code the entry lookups already returned, and treat an unreadable link state as a link so the walk never descends on a guess check the deadline when a directory is popped, not only per entry, so a tree of empty directories cannot outlive the budget read the path parameter once, and reject an invalid limit the way an invalid type is already rejected, instead of silently falling back normalize the resolved path, so a "." or ".." a caller typed reaches neither git nor the client, and return the generic path form with '/' separators on every platform, so the base sent to clients no longer needs a local fixup * ui: expire cached picker searches the cache grew for the lifetime of the component: entries went stale after the TTL but were never removed, so every distinct query typed in a session stayed in memory drop expired entries when a new result is stored * server: address review from @ngxson trim comments to one line each, and drop two that restate the code rename junk_lookup_name to get_effective_name, and move it and the link check to private static members next to junk_dir_names merge the Windows and Linux link checks into one is_link, so symlinks are checked everywhere and junctions only add to it on Windows * server: convert tool paths as UTF-8 on Windows a narrow path uses the active code page there, so a file name came back mangled and a path with an accent could not be opened at all convert explicitly at every crossing between a std::string, which always carries UTF-8 here, and fs::path read the home directory through the wide environment, since the narrow one returns the profile path in the active code page too the walker no longer normalizes separators by hand, since paths now come back in generic form * server: fold the platform branch inside console_output_to_utf8 match the shape of the other helpers, one definition with the #if inside, instead of two definitions wrapped in #if and #else inline the single caller helper and trim the comment
tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034)
llama.cpp
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / dev branches / compile times / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
- subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55.1%
C
16.1%
Python
7.2%
Cuda
5.5%
TypeScript
4.4%
Other
11.5%