Files
ik_llama.cpp/models/templates
Yap Sok AnnandGitHub 6b55d2c750 Fix DSV4 tool calls and reasoning (#2242)
* Fix DSV4 tool calls and reasoning

There are multiple changes. The most important one is the wiring, to
avoid falling back to the autoparser. With autoparser, all arguments
will be forced by the generated grammar to have the `string="true"`
attribute, which then breaks prompt caching, as it would diverge from
what is rendered by the template. Parallel tool calls also doesn't work
when falling back to autoparser.

Other changes:
* Order tool results by tool call order.
* Consume `</think>` instead of `<think></think>` when thinking is
  disabled.
* Use "preserved thinking" mode when any tool is defined, otherwise use
  "interleaved thinking" mode, e.g. for multi-turns chat. Set template
  arg `drop_thinking` to false to force "preserved thinking" mode even
  when no tool is defined.
* Add a message to system prompt when reasoning effort is set to max.

The changes were made by following:
1. The Technical Report: https://arxiv.org/abs/2606.19348
2. Reference implementatin: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py
3. VLLM implementation: https://github.com/vllm-project/vllm/blob/main/vllm/tokenizers/deepseek_v4_encoding.py

For the last bullet point about max reasoning effort, the reference
implementation diverges from the other 2, so we follow the techinical
report and the VLLM implementation, for now. This needs more testing.

* Add back trailing newline

* Update the reasoning effort instruction to follow the reference impl

Using the prompt counting test from @coder543, 0731 does have a special
instruction for "high" and another special instruction for "max".

This will break preview, but assuming most people will use the 0731
release, it should be fine.

[1] https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openrouter_reasoning_effort_levels_are_broken_for/
2026-08-04 19:28:06 +03:00
..

These templates can be updated with the following commands:

./scripts/get_chat_template.py CohereForAI/c4ai-command-r-plus tool_use      > models/templates/CohereForAI-c4ai-command-r-plus-tool_use.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 default  > models/templates/CohereForAI-c4ai-command-r7b-12-2024-default.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 rag      > models/templates/CohereForAI-c4ai-command-r7b-12-2024-rag.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 tool_use > models/templates/CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Llama-8B      > models/templates/deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Qwen-32B      > models/templates/deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja
./scripts/get_chat_template.py fireworks-ai/llama-3-firefunction-v2          > models/templates/fireworks-ai-llama-3-firefunction-v2.jinja
./scripts/get_chat_template.py google/gemma-2-2b-it                          > models/templates/google-gemma-2-2b-it.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.1               > models/templates/meetkai-functionary-medium-v3.1.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.2               > models/templates/meetkai-functionary-medium-v3.2.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.1-8B-Instruct              > models/templates/meta-llama-Llama-3.1-8B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.2-3B-Instruct              > models/templates/meta-llama-Llama-3.2-3B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.3-70B-Instruct             > models/templates/meta-llama-Llama-3.3-70B-Instruct.jinja
./scripts/get_chat_template.py microsoft/Phi-3.5-mini-instruct               > models/templates/microsoft-Phi-3.5-mini-instruct.jinja
./scripts/get_chat_template.py mistralai/Mistral-Nemo-Instruct-2407          > models/templates/mistralai-Mistral-Nemo-Instruct-2407.jinja
./scripts/get_chat_template.py NousResearch/Hermes-2-Pro-Llama-3-8B tool_use > models/templates/NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja
./scripts/get_chat_template.py NousResearch/Hermes-3-Llama-3.1-8B tool_use   > models/templates/NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja
./scripts/get_chat_template.py Qwen/Qwen2.5-7B-Instruct                      > models/templates/Qwen-Qwen2.5-7B-Instruct.jinja
./scripts/get_chat_template.py Qwen/QwQ-32B                                  > models/templates/Qwen-QwQ-32B.jinja
./scripts/get_chat_template.py Qwen/Qwen3-0.6B                               > models/templates/Qwen-Qwen3-0.6B.jinja
./scripts/get_chat_template.py zai-org/GLM-4.5                               > models/templates/zai-org-GLM-4.5.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V3.1                     > models/templates/deepseek-ai-DeepSeek-V3.1.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V4                       > models/templates/deepseek-ai-DeepSeek-V4.jinja