Skip to content
Go back

Local LLM JSON Output That Doesn't Lie

By KingPin 13 min read
Local LLM JSON Output That Doesn't Lie
Contents

The pipeline that parsed fine and wrote nonsense

It’s 2 AM. Your invoice extractor has run for six hours. Zero parse errors. Zero exceptions. Every row landed in Postgres with the right keys and the right types.

Then accounting asks why a $48 receipt is stored as $48,000. The JSON was perfect. The number was invented.

That is the lesson of structured output on local models. Grammar-constrained decoding guarantees the syntax: valid JSON that matches your schema’s shape. It never guarantees the truth of the values inside. A model forced to emit {"total": <number>} will emit a number, and it will pick one even when the receipt has no total on it.

This post compares how llama.cpp (llama-server) and Ollama expose constrained decoding, mentions vLLM for people who serve with it, and ends with the part that saves your 2 AM: validate on the client anyway, and retry. If you want the API-side basics (when JSON mode helps and when it hurts reasoning), that is covered in When to Use Structured Output (JSON Mode) in LLMs. This one is about the local stack and its sharp edges.

How constrained decoding works (a mask, nothing more)

A model produces a score for every token in its vocabulary at each step. Normally the sampler picks from the likely ones. With constrained decoding, a grammar tracks where you are in the output and marks which tokens are legal next. Illegal tokens get their probability set to zero before sampling.

After {"sentiment": the only legal tokens are the opening quote of your allowed strings. The model cannot wander off and write “Sure! Here’s your JSON:”. The mask forbids it.

Two consequences matter:

  1. The model gets no smarter, only fenced in. If it had no idea what the right value was, the fence still forces a value of the right type.
  2. The model usually does not see your schema. In llama.cpp the schema is only used to build the constraint and is not injected into the prompt. The llama.cpp docs say so directly: the model has no visibility into the schema, so describe the structure in your prompt. Ollama’s docs give the same advice and suggest passing the schema as a string in the prompt too.

That second point is the number one cause of “valid JSON, garbage values”.

JSON mode versus schema mode

There are two different features that both get called “JSON mode”:

What it enforcesYour keys
Plain JSON mode (Ollama format: "json", response_format: {"type": "json_object"})Output parses as a JSON objectNot enforced
Schema mode (Ollama format: {...schema...}, llama.cpp json_schema)Output matches the schema’s shapeEnforced

Plain JSON mode gives you {"answer": "yes"} today and {"result": {"verdict": true}} tomorrow. Both parse. Neither matches what your code expects. Use schema mode, always.

llama.cpp: schema in, grammar out

llama-server converts a subset of JSON Schema to a GBNF grammar for you. Start the server:

Terminal window
llama-server -m ./model.gguf --port 8080 -c 8192

On the OpenAI-style chat endpoint, pass the schema inside response_format:

Terminal window
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "system", "content": "Extract the receipt. Reply with JSON: vendor (string), total (number, or null if absent)."},
{"role": "user", "content": "ACME HARDWARE\n2x bolts 3.00\n1x drill 45.00\nthanks!"}
],
"temperature": 0,
"max_tokens": 300,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "receipt",
"schema": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": ["number", "null"]}
},
"required": ["vendor", "total"]
}
}
}
}' | jq -r '.choices[0].message.content'

Note the system prompt repeats the structure in plain words. The grammar does not do that for you.

On the native /completion endpoint the field is called json_schema and sits at the top level of the body. The same endpoint accepts a raw grammar field if you want to write GBNF yourself. That is handy for output that is not JSON at all, like a closed list of labels:

root ::= "{" ws "\"sentiment\"" ws ":" ws ("\"positive\"" | "\"negative\"" | "\"neutral\"") ws "}"
ws ::= " "?
Terminal window
jq -Rn --rawfile g sentiment.gbnf \
'{prompt: "Review: the fan is loud but the cooling is great.\nJSON sentiment:", grammar: $g, n_predict: 64}' \
| curl -s http://localhost:8080/completion -H "Content-Type: application/json" -d @-

That ws rule is deliberate. A classic grammar writing mistake is ws ::= [ \t\n]*, which allows unlimited whitespace. A model that gets stuck can then spend its whole token budget on spaces and newlines. Bound it.

Which schema features llama.cpp ignores

The llama.cpp grammar docs list the converter’s known limits. The important part: unsupported features are skipped silently. Your schema “works”, the constraint is just looser than you think. The documented gaps include:

Also: additionalProperties defaults to false in llama.cpp, the opposite of the JSON Schema spec, on purpose (faster grammars, fewer hallucinated keys). If you rely on extra keys, set it to true explicitly.

So test your schema before you trust it: feed it a few deliberately wrong outputs and confirm the server rejects them, or check strings against the generated grammar with the test-gbnf-validator binary that a llama.cpp build with tests produces. Do this once per schema, not once per outage.

Ollama: format takes a string or a schema

Ollama’s chat and generate endpoints have one field, format. Set it to the string "json" and you get plain JSON mode. Set it to a JSON schema object and you get schema mode.

Terminal window
curl -s http://localhost:11434/api/chat -d '{
"model": "qwen3.6:35b",
"stream": false,
"messages": [
{"role": "user", "content": "Extract the receipt as JSON with vendor and total (number or null):\nACME HARDWARE\n2x bolts 3.00\n1x drill 45.00"}
],
"options": {"temperature": 0, "num_predict": 300},
"format": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": ["number", "null"]}
},
"required": ["vendor", "total"]
}
}' | jq -r '.message.content'

qwen3.6:35b and gemma4:12b are both current tags in the Ollama library as of October 2026. Any model you already use works the same way, since the constraint lives in the runtime and not in the model.

Ollama’s own API docs carry a warning for plain JSON mode: tell the model to answer in JSON in the prompt, or it may emit large amounts of whitespace. That is the whitespace loop from the grammar section, showing up in a different place. Schema mode reduces the risk, and a prompt that names the structure removes most of the rest.

Gotchas that bite after the first demo

The model fills fields with confident filler

You gave it a schema it cannot see. It knows vendor is a string. It does not know you meant the store name and not the street address. Describe each field in the prompt, and keep description strings in the schema as a second copy for humans. For Ollama, you can also paste the schema as text into the prompt, as their docs suggest.

Property order is reasoning order

A model writes tokens left to right. If your schema lists answer before reasoning, the model commits to the answer first and then writes a justification for whatever it already said. Put the thinking field first:

{
"type": "object",
"properties": {
"reasoning": {"type": "string"},
"total": {"type": ["number", "null"]}
},
"required": ["reasoning", "total"]
}

Pydantic emits properties in field declaration order, and the grammar llama.cpp generates follows the schema’s order (you can see it in the sample grammar in their docs). On Ollama, check a few real outputs for your model rather than assuming. This is a cheap, free accuracy bump for small models.

Enums stop invention

If a field has a closed set of values, say so with enum. Without it, a model will give you "urgent", "Urgent", "URGENT!", and eventually "kinda urgent". With it, the mask leaves only the legal strings.

Truncation leaves invalid JSON

The grammar guarantees every prefix is valid. It cannot guarantee the output finishes. If generation hits the token limit (max_tokens or n_predict on llama-server, num_predict under Ollama options) halfway through an array, you get cut-off JSON that will not parse. Check the stop reason, which the OpenAI-style response reports as finish_reason (stop_type on llama-server’s native /completion) and Ollama reports as done_reason. Anything other than a normal stop means “do not trust this”. Set the limit well above your longest expected output.

Thinking models and constraints

Reasoning models that emit a thinking block add a wrinkle. llama-server accepts chat_template_kwargs such as {"enable_thinking": false} in the request, which turns off thinking for templates that support it. Test with thinking on and off on your schema. Which behavior you get depends on the model template and server version, so treat your own test run as the authority.

Python: Pydantic as the single source of truth

Write the schema once as a Pydantic model. Generate the JSON schema from it. Validate the reply with it. Here it is against Ollama:

extract.py
from ollama import chat
from pydantic import BaseModel, model_validator
class Receipt(BaseModel):
reasoning: str
vendor: str
total: float | None
@model_validator(mode="after")
def total_is_sane(self):
if self.total is not None and not (0 < self.total < 100_000):
raise ValueError("total out of range")
return self
def extract(text: str) -> Receipt:
resp = chat(
model="qwen3.6:35b",
messages=[
{
"role": "system",
"content": "Extract the receipt as JSON: reasoning (short), vendor, "
"total (number, or null if no total is printed).",
},
{"role": "user", "content": text},
],
format=Receipt.model_json_schema(),
options={"temperature": 0, "num_predict": 400},
)
return Receipt.model_validate_json(resp.message.content)

The ollama Python library accepts a schema dict for format, and model_json_schema() produces one. Their docs use this exact pattern.

Against llama-server you use the standard openai client, since the server speaks the OpenAI chat API. This reuses Receipt from extract.py, so either keep the ollama package installed or copy the class into its own module:

extract_llamacpp.py
from openai import OpenAI
from extract import Receipt
client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="local", # llama-server serves the model it loaded
messages=[
{"role": "system", "content": "Extract the receipt as JSON: reasoning, vendor, total."},
{"role": "user", "content": "ACME HARDWARE\n1x drill 45.00"},
],
temperature=0,
max_tokens=400,
response_format={
"type": "json_schema",
"json_schema": {"name": "receipt", "schema": Receipt.model_json_schema()},
},
)
receipt = Receipt.model_validate_json(resp.choices[0].message.content)

One Pydantic caveat: models with nested sub-models make model_json_schema() emit $defs and $ref. Given the nested $ref limits above, flatten deep models for llama.cpp, or test the generated grammar.

Validate and retry, because the mask is not a judge

Here is the loop that would have caught the $48,000 receipt. The validator above rejects absurd totals. When validation fails, feed the error back and try again:

retry.py
from ollama import chat
from pydantic import ValidationError
from extract import Receipt
def extract_with_retry(text: str, attempts: int = 3) -> Receipt:
messages = [
{"role": "system", "content": "Extract the receipt as JSON: reasoning, vendor, total or null."},
{"role": "user", "content": text},
]
last_error = None
for _ in range(attempts):
resp = chat(
model="qwen3.6:35b",
messages=messages,
format=Receipt.model_json_schema(),
options={"temperature": 0.3, "num_predict": 400},
)
raw = resp.message.content
try:
return Receipt.model_validate_json(raw)
except ValidationError as e:
last_error = e
messages += [
{"role": "assistant", "content": raw},
{"role": "user", "content": f"That failed validation: {e}. Fix it and reply with JSON only."},
]
raise RuntimeError(f"gave up after {attempts} tries: {last_error}")

Two details. First, retries at temperature 0 tend to repeat themselves, so nudge it up a little. Second, a retry fixes bad values only when your validators can recognize them. A schema alone cannot tell a right total from a wrong total. Add checks that tie fields to each other: totals equal the sum of line items, dates fall in the range you expect, IDs exist in your database. Those checks are where the truth gets enforced.

vLLM, briefly

If you serve with vLLM, structured outputs are built into the OpenAI-compatible server. vLLM docs name xgrammar and guidance as backends, and the --structured-outputs-config.backend flag to vllm serve selects one, with auto as the default that picks a backend per request. The standard response_format with json_schema works as shown above. vLLM’s own fields go inside extra_body={"structured_outputs": {"json": schema}}. The older guided_json family of fields was removed in v0.12.0, so tutorials from 2024 will fail on a current install.

Same rules apply: the syntax is guaranteed, the values are not.

The SumGuy Take

Use schema mode, not plain JSON mode. Put the structure in the prompt too. Put reasoning before the answer. Use enums for closed sets. Leave generous token headroom and check the stop reason. Then run every reply through Pydantic with validators that know something about your data, and retry with the error message.

Constrained decoding is a seatbelt. It keeps the output on the road. It does not drive.

Common Questions

Does Ollama structured output work with every model?

Yes for text models in practice, because the constraint is applied by the Ollama runtime during sampling, not by the model itself. Quality still varies: small models follow your field descriptions worse, so they fill valid fields with weaker values. Test your schema on the exact model tag you deploy.

Why does my local model return valid JSON with wrong values?

Grammar-constrained decoding only restricts which tokens are legal, so the model must emit something of the right type even when it lacks the information. Fix it with field descriptions in the prompt, nullable fields for missing data, enums for closed sets, and Pydantic validators that cross-check values.

Does llama.cpp validate my JSON Schema?

No. The llama.cpp schema converter skips unsupported keywords silently, so a schema can load without errors and still constrain less than you expect. Test the generated grammar with the test-gbnf-validator test binary, and validate replies on the client with Pydantic.

Should I use GBNF grammars or JSON schema in llama-server?

Use JSON schema for JSON output, since llama-server converts it to GBNF for you. Write raw GBNF when the output is not JSON, such as a closed list of labels or a custom line format. Bound whitespace rules in handwritten grammars so a stuck model cannot loop on spaces.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
AppArmor Profiles for Docker, Step by Step

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts