mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-12 22:31:11 +04:00
The reference implementation has two mutually exclusive prompt layouts. In non streaming mode the prefill carries the whole utterance text plus tts_eos summed with codec_pad, and the trailing text hidden collapses to a single tts_pad row. In streaming mode the prefill carries only the first text token and the trailing rows stream the rest of the text followed by tts_eos. The pipeline built the non streaming prefill but the streaming overlay, so the talker saw the utterance a second time during generation and read it twice before emitting codec_eos. The overlay is now the single tts_pad row that matches the prefill.