Files
ik_llama.cpp/examples
e5357286c0 hotswap: keep load-time-derived and transformed tensors coherent after reloads (follow-up to #2131) (#2163)
* hotswap: keep load-time-derived and transformed tensors coherent after reloads

- Re-derive the MLA combined attn_kv_b (computed_wkv_b) in place when a layer's
  attn_k_b/attn_v_b are hot-swapped: the mla>1 prompt-processing path consumes
  the derived tensor, so swaps of the source tensors previously had no effect
  (KLD stayed exactly 0 in per-tensor benchmarks, e.g. GLM-5.2 attn_k_b/v_b).
- Refuse (loudly) hot-swaps that cannot be correct: views into -mqkv/-muge
  merged tensors, khad-folded MLA weights, in-place-scaled ffn_gate_inp_s,
  BitNet fused scales, OpenPangu parameter-sink sources, and same-dtype swaps
  of mmap-backed tensors. A refused reload produces no 'reloaded tensor' line,
  so benchmark drivers quarantine the round instead of recording wrong data.
- Propagate reloaded data to same-name duplicate instances (tied lm head copy
  of token_embd, per-layer rope_freqs/rope_factors copies, expert-bias dups),
  warning when a duplicate cannot be refreshed.
- Warn that derived state stays stale where a refresh is not possible:
  pre-transposed wk_b_pp under -sm graph/attn, requantized MTP head
  (output_extra.weight), and k_b/v_b derived from a reloaded attn_kv_b.
- Warn at registration time when -rtr is enabled (restores cannot reproduce
  the run-time-repacked state; F16 -> BF16_R16 is lossy).
- server: only attempt the /health hot-swap reload when no slot is processing,
  and clear the KV cache + cached prompts after a successful reload (they were
  computed with the previous weights).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Deduplicate llm_compute_wkv_b

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 17:34:43 +03:00
..
2026-04-23 09:05:39 +02:00
2024-07-27 07:55:01 +02:00
2024-07-27 07:55:01 +02:00
2024-08-12 15:14:32 +02:00
2023-03-29 20:21:09 +03:00
2024-07-27 07:55:01 +02:00