mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-12 22:31:11 +04:00
* Get started with Onyx
* Add architecture
* Skip keys handled in super()
* Loading tensors
* Shorten
* Graph
* Apply suggestion from @pcuenca
* Remove norm now embedding in transformers weights
* Add eot
* Explicit output_multiplier
* Handle post_norm_eps
* No super call; unhardcode eot.
The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.
* Register for drafting
* DFlash: inherit rope type from the linked target.
Another option would be to store it in the gguf file itself.
* mmproj conversion
Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.
* "clip" header declarations
* Load mmproj
* Pre-processing
* Graph
* Go back to using delimiters.
Otherwise our generations are worse.
Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.
* downsample_factor -> merge_size
* Add vision graph
lol, forgot from a previous commit
* Additional renames, align with llama.cpp / transformers
* Prefer _size instead of independent _h and _w
* Fix token layout
Co-authored-by: Young Han <younghan@fb.com>
* onyx: bring the chat parser onto the onyx branch
common/chat.cpp on this branch has no Onyx handling, so a converted model
serves malformed chat: the assistant preamble leaks into content
("to=self<|message|>...") and tool calls fail with
HTTP 500 "The model produced output that does not match the expected
peg-native format"
common_chat_params_init_onyx exists on onyx-fair-patch, added there by
8bb73dd3d. It was never on this branch, so this is not a regression --
the two lines developed independently.
The code here is taken verbatim from that commit. It is the clean side of
`git merge origin/onyx-fair-patch`: chat.cpp is one of the files that
merges without conflict. The full merge is not viable -- it produces 13
conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp
where the q_norm-folding and metadata-scale approaches contradict each
other, and #4/#7 are stacked on this branch's side of that.
Verified on this branch: builds with 0 errors, converts an Onyx checkpoint,
and serving it gives "4" for "What is 2+2?" plus a correct
get_weather {"city":"Paris"} tool call, where the unported branch gives the
two failures above.
No converter or runtime changes are included, so this should not interact
with the q_norm work.
Co-authored-by: Beto de Paola <betodepaola@meta.com>
* Less params, bilinear pos-emb interpolation as a graph op instead of CPU
* Map to symbolic V_MMPROJ instead of strings
* Make a couple params explicit
* Patchify via build_inp()
* No param for rope_theta
* Small cleanup
* Restore blank line
* Unpermute, to adapt to the latest transformers checkpoint
* Apply norm after token embeddings
This follows the latest transformers approach.
* Remove duplicated function
* build_vit
* onyx: use the model rope theta on sliding-window layers
* DFlash: conversion from transformers drafter
* Revert rope_type derivation from target
NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as
the Q/K are stored in "NEOX" (rotated half) format, like in
transformers.
* Apply suggestion from @pcuenca
* Set model type
* Remove comment that will become obsolete
* Hardcode post_norm_rms_eps instead of new param
* Derive SWA+RoPE pattern from gguf array or scalar
* Fix model type <-> number of layers
* Reorder
* Rename
* Fix typo
* DFlash: seed the draft KV cache from multimodal embedding batches
`common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch:
```
decoding image batch 1/1, n_tokens_batch = 256
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0)
srv decode: failed to process speculative batch
```
Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through.
Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix.
Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request:
- before: HTTP 500, `failed to process speculative batch`
- after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04
Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing.
* Conversion: prefer rewrite to mapping
* Revert "Conversion: prefer rewrite to mapping"
This reverts commit a92d0ac584.
* fix lint
* sliding_window metadata is not optional
* disable state save/load
* Apply suggestion from @pcuenca
---------
Co-authored-by: Young Han <younghan@fb.com>
Co-authored-by: Beto de Paola <betodepaola@meta.com>
Co-authored-by: Daniel Han <michaelhan2050@gmail.com>
Co-authored-by: ruanrms <ruanslv@gmail.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
180 lines
7.7 KiB
Python
180 lines
7.7 KiB
Python
from __future__ import annotations
|
|
|
|
import json
|
|
from typing import Any, Iterable, TYPE_CHECKING
|
|
|
|
import torch
|
|
|
|
if TYPE_CHECKING:
|
|
from torch import Tensor
|
|
|
|
from .base import MmprojModel, ModelBase, TextModel, gguf
|
|
|
|
|
|
def _unpermute_for_rope(tensor: "Tensor", n_heads: int) -> "Tensor":
|
|
"""Invert transformers' `_permute_for_rope`: HF stores Q/K in rotate_half layout,
|
|
llama.cpp consumes the interleaved (NORM) layout."""
|
|
if tensor.ndim == 2:
|
|
dim1, dim2 = tensor.shape
|
|
return tensor.view(n_heads, 2, dim1 // n_heads // 2, dim2).transpose(1, 2).reshape(dim1, dim2)
|
|
if tensor.ndim == 1:
|
|
(dim1,) = tensor.shape
|
|
return tensor.view(n_heads, 2, dim1 // n_heads // 2).transpose(1, 2).reshape(dim1)
|
|
raise ValueError(f"_unpermute_for_rope: unexpected shape {tuple(tensor.shape)}")
|
|
|
|
|
|
@ModelBase.register("MuseGlimmerForConditionalGeneration")
|
|
class MuseGlimmerModel(TextModel):
|
|
model_arch = gguf.MODEL_ARCH.MUSE_GLIMMER
|
|
|
|
def norm_shift(self, name: str) -> float:
|
|
# All four layer norms use 1, the final norm uses 0.
|
|
return 1.0 if name.endswith("layernorm.weight") else 0.0
|
|
|
|
def set_vocab(self):
|
|
self._set_vocab_gpt2()
|
|
|
|
from transformers import AutoTokenizer
|
|
tok = AutoTokenizer.from_pretrained(self.dir_model)
|
|
eot_id = tok.convert_tokens_to_ids("<|eot|>")
|
|
if isinstance(eot_id, int) and eot_id >= 0:
|
|
self.gguf_writer.add_eot_token_id(eot_id)
|
|
|
|
def set_gguf_parameters(self):
|
|
super().set_gguf_parameters()
|
|
hparams = self.hparams
|
|
|
|
self.gguf_writer.add_final_logit_softcapping(hparams["final_logit_softcapping"])
|
|
self.gguf_writer.add_logit_scale(hparams["output_multiplier"])
|
|
self.gguf_writer.add_sliding_window(hparams["sliding_window"])
|
|
self.gguf_writer.add_sliding_window_pattern([t == "sliding_attention" for t in hparams["layer_types"]])
|
|
|
|
def modify_tensors(self, data_torch: Tensor, name: str, bid: int | None) -> Iterable[tuple[str, Tensor]]:
|
|
shift = self.norm_shift(name)
|
|
if shift != 0.0:
|
|
data_torch = data_torch + shift
|
|
|
|
# Invert transformers' `_permute_for_rope` on Q/K, we keep ggml's NORM (interleaved) rope
|
|
if ".self_attn.q_proj." in name:
|
|
data_torch = _unpermute_for_rope(data_torch, int(self.hparams["num_attention_heads"]))
|
|
elif ".self_attn.k_proj." in name:
|
|
data_torch = _unpermute_for_rope(data_torch, int(self.hparams["num_key_value_heads"]))
|
|
|
|
# Synthesize QK-norm weights to absorb qk_scale_factor.
|
|
# MuseGlimmer implementation: scaleless RMSNorm followed by qk_scale_factor..
|
|
if bid is not None and name.endswith(f"model.layers.{bid}.self_attn.q_proj.weight"):
|
|
head_dim = self.hparams["head_dim"]
|
|
q_scale = float(self.hparams["qk_scale_factor"])
|
|
yield (
|
|
self.map_tensor_name(f"model.layers.{bid}.self_attn.q_norm.weight"),
|
|
torch.full((head_dim,), q_scale, dtype=torch.float32),
|
|
)
|
|
yield (
|
|
self.map_tensor_name(f"model.layers.{bid}.self_attn.k_norm.weight"),
|
|
torch.ones((head_dim,), dtype=torch.float32),
|
|
)
|
|
|
|
yield from super().modify_tensors(data_torch, name, bid)
|
|
|
|
|
|
@ModelBase.register("MuseGlimmerForConditionalGeneration")
|
|
class MuseGlimmerVisionModel(MmprojModel):
|
|
def get_vision_config(self) -> dict[str, Any] | None:
|
|
c = self.global_config.get("vision_config")
|
|
if not c:
|
|
return None
|
|
# MuseGlimmer actually uses dynamic size, initialize with nominal size
|
|
image_size = c["pos_emb_height"] * c["patch_size"] * c["merge_size"]
|
|
return {**c, "image_size": image_size}
|
|
|
|
def set_gguf_parameters(self):
|
|
super().set_gguf_parameters()
|
|
assert self.hparams_vision is not None
|
|
c = self.hparams_vision # enriched vision_config from get_vision_config()
|
|
|
|
self.gguf_writer.add_clip_projector_type(gguf.VisionProjectorType.MUSE_GLIMMER)
|
|
self.gguf_writer.add_vision_attention_layernorm_eps(float(c["layer_norm_eps"]))
|
|
self.gguf_writer.add_vision_spatial_merge_size(int(c["merge_size"]))
|
|
|
|
@classmethod
|
|
def filter_tensors(cls, item):
|
|
name, gen = item
|
|
keep = ("model.vision_tower.", "model.vision_adapter.", "model.vision_projection.")
|
|
if not any(name.startswith(k) for k in keep):
|
|
return None
|
|
return super().filter_tensors((name, gen))
|
|
|
|
# 3-layer projector MLP
|
|
_MM_MLP_MAP = {
|
|
"model.vision_adapter.fc1": (gguf.MODEL_TENSOR.V_MMPROJ, 0),
|
|
"model.vision_adapter.fc2": (gguf.MODEL_TENSOR.V_MMPROJ, 1),
|
|
"model.vision_projection": (gguf.MODEL_TENSOR.V_MMPROJ, 2),
|
|
}
|
|
|
|
def modify_tensors(self, data_torch, name, bid):
|
|
assert self.hparams_vision is not None
|
|
if ".attn.q_proj." in name or ".attn.k_proj." in name:
|
|
n_heads = int(self.hparams_vision["num_attention_heads"])
|
|
data_torch = _unpermute_for_rope(data_torch, n_heads)
|
|
# Lay out the pt=2 temporal slabs of the patch embedding as a conv2d for build_inp()
|
|
if name.endswith("patch_embedder.patch_embedding.weight"):
|
|
n_embd = data_torch.shape[0]
|
|
pt = int(self.hparams_vision["patch_temporal"])
|
|
ps = int(self.hparams_vision["patch_size"])
|
|
data_torch = data_torch.view(n_embd, pt, 3, ps, ps).sum(dim=1) # (n_embd, 3, ps, ps)
|
|
stem, _, suffix = name.rpartition(".")
|
|
if stem in self._MM_MLP_MAP:
|
|
tensor_key, idx = self._MM_MLP_MAP[stem]
|
|
yield (self.format_tensor_name(tensor_key, bid=idx, suffix="." + suffix), data_torch)
|
|
return
|
|
yield (self.map_tensor_name(name), data_torch)
|
|
|
|
|
|
@ModelBase.register("MuseGlimmerAssistantModel")
|
|
class MuseGlimmerAssistantModel(TextModel):
|
|
model_arch = gguf.MODEL_ARCH.DFLASH
|
|
|
|
def set_vocab(self):
|
|
if self.target_model_dir is None:
|
|
raise ValueError(
|
|
"MuseGlimmerAssistant (DFlash drafter) requires --target-model-dir pointing to the "
|
|
"target MuseGlimmer HF directory"
|
|
)
|
|
|
|
original_dir = self.dir_model
|
|
self.dir_model = self.target_model_dir
|
|
|
|
from . import get_model_class
|
|
with open(self.target_model_dir / "config.json", "r", encoding="utf-8") as f:
|
|
target_arch = json.load(f)["architectures"][0]
|
|
target_cls = get_model_class(target_arch)
|
|
if target_cls is not type(self):
|
|
target_cls.set_vocab(self) # ty: ignore[unresolved-attribute]
|
|
else:
|
|
super().set_vocab()
|
|
|
|
self.dir_model = original_dir
|
|
|
|
mask_token_id = self.hparams.get("mask_token_id")
|
|
if mask_token_id is not None:
|
|
self.gguf_writer.add_mask_token_id(int(mask_token_id))
|
|
|
|
def set_gguf_parameters(self):
|
|
super().set_gguf_parameters()
|
|
h = self.hparams
|
|
|
|
self.gguf_writer.add_block_size(int(h["block_size"]))
|
|
|
|
# dflash.target_layers[k] refers to the inputs going into the ith layer, which come from the (i-1)th layer's output.
|
|
# The transformers configuration refers to the outputs being recorded.
|
|
self.gguf_writer.add_target_layers([int(x) + 1 for x in h["target_layer_ids"]])
|
|
|
|
if h.get("sliding_window") and h.get("layer_types"):
|
|
self.gguf_writer.add_sliding_window(int(h["sliding_window"]))
|
|
self.gguf_writer.add_sliding_window_pattern([t == "sliding_attention" for t in h["layer_types"]])
|
|
|
|
def modify_tensors(self, data_torch: Tensor, name: str, bid: int | None) -> Iterable[tuple[str, Tensor]]:
|
|
# DFlash defaults to NEOX (rotate_half) rope, matching transformers HF layout for Q/K, QK-norms
|
|
# no permutation needed.
|
|
yield (self.map_tensor_name(name), data_torch)
|