mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-12 22:31:11 +04:00
* Get started with Onyx
* Add architecture
* Skip keys handled in super()
* Loading tensors
* Shorten
* Graph
* Apply suggestion from @pcuenca
* Remove norm now embedding in transformers weights
* Add eot
* Explicit output_multiplier
* Handle post_norm_eps
* No super call; unhardcode eot.
The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.
* Register for drafting
* DFlash: inherit rope type from the linked target.
Another option would be to store it in the gguf file itself.
* mmproj conversion
Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.
* "clip" header declarations
* Load mmproj
* Pre-processing
* Graph
* Go back to using delimiters.
Otherwise our generations are worse.
Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.
* downsample_factor -> merge_size
* Add vision graph
lol, forgot from a previous commit
* Additional renames, align with llama.cpp / transformers
* Prefer _size instead of independent _h and _w
* Fix token layout
Co-authored-by: Young Han <younghan@fb.com>
* onyx: bring the chat parser onto the onyx branch
common/chat.cpp on this branch has no Onyx handling, so a converted model
serves malformed chat: the assistant preamble leaks into content
("to=self<|message|>...") and tool calls fail with
HTTP 500 "The model produced output that does not match the expected
peg-native format"
common_chat_params_init_onyx exists on onyx-fair-patch, added there by
8bb73dd3d. It was never on this branch, so this is not a regression --
the two lines developed independently.
The code here is taken verbatim from that commit. It is the clean side of
`git merge origin/onyx-fair-patch`: chat.cpp is one of the files that
merges without conflict. The full merge is not viable -- it produces 13
conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp
where the q_norm-folding and metadata-scale approaches contradict each
other, and #4/#7 are stacked on this branch's side of that.
Verified on this branch: builds with 0 errors, converts an Onyx checkpoint,
and serving it gives "4" for "What is 2+2?" plus a correct
get_weather {"city":"Paris"} tool call, where the unported branch gives the
two failures above.
No converter or runtime changes are included, so this should not interact
with the q_norm work.
Co-authored-by: Beto de Paola <betodepaola@meta.com>
* Less params, bilinear pos-emb interpolation as a graph op instead of CPU
* Map to symbolic V_MMPROJ instead of strings
* Make a couple params explicit
* Patchify via build_inp()
* No param for rope_theta
* Small cleanup
* Restore blank line
* Unpermute, to adapt to the latest transformers checkpoint
* Apply norm after token embeddings
This follows the latest transformers approach.
* Remove duplicated function
* build_vit
* onyx: use the model rope theta on sliding-window layers
* DFlash: conversion from transformers drafter
* Revert rope_type derivation from target
NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as
the Q/K are stored in "NEOX" (rotated half) format, like in
transformers.
* Apply suggestion from @pcuenca
* Set model type
* Remove comment that will become obsolete
* Hardcode post_norm_rms_eps instead of new param
* Derive SWA+RoPE pattern from gguf array or scalar
* Fix model type <-> number of layers
* Reorder
* Rename
* Fix typo
* DFlash: seed the draft KV cache from multimodal embedding batches
`common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch:
```
decoding image batch 1/1, n_tokens_batch = 256
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0)
srv decode: failed to process speculative batch
```
Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through.
Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix.
Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request:
- before: HTTP 500, `failed to process speculative batch`
- after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04
Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing.
* Conversion: prefer rewrite to mapping
* Revert "Conversion: prefer rewrite to mapping"
This reverts commit a92d0ac584.
* fix lint
* sliding_window metadata is not optional
* disable state save/load
* Apply suggestion from @pcuenca
---------
Co-authored-by: Young Han <younghan@fb.com>
Co-authored-by: Beto de Paola <betodepaola@meta.com>
Co-authored-by: Daniel Han <michaelhan2050@gmail.com>
Co-authored-by: ruanrms <ruanslv@gmail.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
89 lines
4.1 KiB
C++
89 lines
4.1 KiB
C++
#include "models.h"
|
|
|
|
// MuseGlimmer vision encoder: 50-layer ViT with 2D RoPE, sparse block-diagonal
|
|
// window attention (every 4th + last layer global), pixel-shuffle downsample, then
|
|
// adapter MLP + LLM's vision_projection.
|
|
//
|
|
// Several quantities are precomputed on host and fed as named graph inputs (filled in
|
|
// clip.cpp set_input, PROJECTOR_TYPE_MUSE_GLIMMER branch):
|
|
// muse_glimmer_pos_w/_h [n_tok] i32 : 1-indexed RoPE positions (sparse-permuted order)
|
|
// muse_glimmer_sp_perm [n_tok] i32 : window grouping permutation (applied after ln_pre)
|
|
// muse_glimmer_inv_perm [n_tok] i32 : inverse of sp_perm (applied after blocks)
|
|
// muse_glimmer_ds_perm [n_tok] i32 : pixel-shuffle gather (original order)
|
|
// muse_glimmer_sp_mask [n_tok, n_tok] f32 : block-diagonal window mask (sparse layers)
|
|
ggml_cgraph * clip_graph_muse_glimmer::build() {
|
|
const int ds = hparams.n_merge; // downsample factor (2)
|
|
const int sf = hparams.muse_glimmer_sparse_factor; // 4
|
|
const int n_tok = n_patches;
|
|
const int n_out = (n_patches_x / ds) * (n_patches_y / ds);
|
|
const float rope_base = hparams.rope_theta; // 10000
|
|
|
|
auto inp_i32 = [&](const char * name, int64_t n) {
|
|
ggml_tensor * t = ggml_new_tensor_1d(ctx0, GGML_TYPE_I32, n);
|
|
ggml_set_name(t, name);
|
|
ggml_set_input(t);
|
|
return t;
|
|
};
|
|
|
|
ggml_tensor * pos_w = inp_i32("muse_glimmer_pos_w", n_tok);
|
|
ggml_tensor * pos_h = inp_i32("muse_glimmer_pos_h", n_tok);
|
|
ggml_tensor * sp_perm = inp_i32("muse_glimmer_sp_perm", n_tok);
|
|
ggml_tensor * inv_perm = inp_i32("muse_glimmer_inv_perm", n_tok);
|
|
ggml_tensor * ds_perm = inp_i32("muse_glimmer_ds_perm", n_tok);
|
|
|
|
ggml_tensor * sp_mask = ggml_new_tensor_2d(ctx0, GGML_TYPE_F32, n_tok, n_tok);
|
|
ggml_set_name(sp_mask, "muse_glimmer_sp_mask");
|
|
ggml_set_input(sp_mask);
|
|
|
|
// patchify via build_inp (conv2d over raw pixels) + bilinear-resized learned pos-emb
|
|
ggml_tensor * x = build_inp(); // [n_embd, n_tok, 1]
|
|
x = ggml_add(ctx0, x, resize_position_embeddings(GGML_SCALE_MODE_BILINEAR));
|
|
cb(x, "after_posemb", -1);
|
|
|
|
// group patches into pgrid x pgrid windows (sparse attention order)
|
|
x = ggml_get_rows(ctx0, x, sp_perm);
|
|
cb(x, "after_sp_perm", -1);
|
|
|
|
// per-layer mask: sparse layers get sp_mask, global layers (every sf-th and last) get none
|
|
std::vector<ggml_tensor *> attn_mask_layers(n_layer);
|
|
for (int il = 0; il < n_layer; ++il) {
|
|
const bool is_global = (il == n_layer - 1) || ((il + 1) % sf == 0);
|
|
attn_mask_layers[il] = is_global ? nullptr : sp_mask;
|
|
}
|
|
|
|
// 2D RoPE: first half of head_dim uses width pos, second half uses height pos
|
|
auto add_pos = [&](ggml_tensor * cur, const clip_layer &) {
|
|
return build_rope_2d(ctx0, cur, pos_w, pos_h, rope_base, false);
|
|
};
|
|
|
|
build_vit_opts opts;
|
|
opts.attn_mask_layers = std::move(attn_mask_layers);
|
|
|
|
// pre_ln, per-layer transformer, post_ln (all inside build_vit); reference uses exact (erf) GELU
|
|
x = build_vit(x, n_tok, NORM_TYPE_NORMAL, FFN_GELU_ERF, nullptr, add_pos, opts);
|
|
|
|
// un-permute back to original grid order
|
|
x = ggml_get_rows(ctx0, x, inv_perm);
|
|
cb(x, "after_inv_perm", -1);
|
|
|
|
// pixel-shuffle downsample: gather f*f spatial neighbors then concat channel-outer.
|
|
// out[c*(ds*ds)+s, o] = x[ds_perm gathered][o*(ds*ds)+s, c]
|
|
x = ggml_get_rows(ctx0, x, ds_perm); // [n_embd, n_tok], grouped
|
|
x = ggml_reshape_3d(ctx0, x, n_embd, ds * ds, n_out);// [c, s, o]
|
|
x = ggml_permute(ctx0, x, 1, 0, 2, 3); // [s, c, o]
|
|
x = ggml_cont(ctx0, x);
|
|
x = ggml_reshape_2d(ctx0, x, n_embd * ds * ds, n_out); // [6144, n_out]
|
|
cb(x, "encoder_out", -1);
|
|
|
|
// adapter (6144->4096->4096, exact GELU each) + LLM vision_projection (4096->6656)
|
|
x = build_mm(model.mm_0_w, x);
|
|
x = ggml_gelu_erf(ctx0, x);
|
|
x = build_mm(model.mm_1_w, x);
|
|
x = ggml_gelu_erf(ctx0, x);
|
|
x = build_mm(model.mm_2_w, x); // [6656, n_out]
|
|
cb(x, "projected", -1);
|
|
|
|
ggml_build_forward_expand(gf, x);
|
|
return gf;
|
|
}
|