Local model chat templates convert a conversation into the token sequence a model expects. A downloaded model can load successfully and still receive poorly formatted chat input. The role labels, control tokens and start of the next assistant turn matter alongside the words in your question.
A template is therefore part of using a chat model, rather than an optional decorative wrapper. It tells the runtime how to arrange messages for a model trained to recognize a particular conversation format.
A conversation becomes one sequence
The Hugging Face Transformers guide explains that chat models still continue a sequence of tokens. A chat interface represents the conversation as messages, commonly with role and content fields. The template turns those messages into the format passed to the model.
The guide illustrates two instruction models derived from the same base model that use different control tokens. Sharing a base model does not establish that two fine-tuned models share a chat format.
This distinction is useful when comparing local downloads. The weights determine the model being run, while the tokenizer and template help determine how your conversation reaches those weights. A filename saying “Instruct” is not enough information to recreate the format yourself.
Roles need the model’s own delimiters
The familiar roles user, assistant and system describe who contributes a message. Their exact serialized representation can vary between models.
A word such as “user” printed into ordinary text is not necessarily equivalent to the special token or delimiter used by the model’s template. Adding homemade role headings can therefore produce a sequence different from the one the model expects.
For a new download, start with the tokenizer and template supplied for that model. When investigating unexpected behavior, preserve a small conversation and inspect the formatted result before changing several other settings at once.
That is a troubleshooting recommendation based on the documented formatting mechanism. It does not imply that every poor response has a template error.
Starting a new assistant turn and continuing one differ
The Transformers API exposes add_generation_prompt to append the tokens indicating the start of an assistant response. The exact effect depends on the template; some model formats do not need extra start tokens.
It also documents continue_final_message for continuing the final message rather than opening another turn. This is relevant to a prefill: an application can provide the beginning of the assistant’s answer and ask the model to continue it.
Those options express different intentions. One begins a reply; the other extends text already supplied. The guide says they should not be used together.
| Intention | Formatting question |
|---|---|
| Answer the user’s latest message | Does the format mark the start of the assistant’s turn? |
| Continue a supplied assistant prefill | Does the final message remain open for continuation? |
| Include an earlier assistant response | Is that message closed before the following turn? |
When a local model starts extending a user’s question, inspect the boundary at the end of the input. A missing or inappropriate turn marker is one possibility to investigate.
Avoid adding special tokens twice
A template can already include the special tokens needed by the model. If an application first formats the chat as text and then tokenizes that text, it can accidentally add another set of special tokens.
Transformers documents add_special_tokens=False when tokenizing text produced with apply_chat_template(tokenize=False). Using the template’s tokenizing path avoids that extra formatting step.
The point is not that special tokens are undesirable. They belong in the sequence once, in the positions expected by the selected model. A duplicated beginning marker is a formatting change even though the visible conversation looks the same.
Keep the formatted text and the tokenization step in the same review. Checking only the message list hides what the runtime actually submitted.
Check what the local runtime reads
llama.cpp’s server documentation documents chat-template handling and Jinja support. Options and supported behavior can change with the installed build, so match instructions to that build’s documentation and help output.
The GGUF guide explains the role of a local model file and its metadata. A runtime’s treatment of the template matters in addition to the file’s presence.
Do not assume that a server exposing a familiar chat API has reproduced every behavior of a hosted service. Matching request field names is a different question from matching the model, template and tool-handling behavior.
Investigate one layer at a time
Begin with a minimal conversation, the model’s intended tokenizer and its supplied template. Inspect the serialized chat, then check the start or continuation of the next assistant message.
If the format is correct, continue to other causes: the model variant, task suitability, generation settings and the surrounding application. Changing the template to hide an unrelated limitation can make later comparisons harder to interpret.
Memory is another separate constraint. The KV cache guide follows the context-related memory that can grow after weights have loaded. A correct template cannot create additional memory, although a conversation’s length affects the runtime’s workload.
Keep a short record of the model revision, runtime version and formatting settings when reporting a problem. Those details help another person reproduce the request without claiming a diagnosis from the response alone.




