My 8GB GPU shouldn't run flagship local LLMs, but this workflow makes it work anyway
XDA ·
Follow into
Save into
Follow into

Local LLM releases seem to keep getting bigger, and many of the ones people actually get excited about now need a lot more memory than my card has. That's great news if you've got a 24+ GB GPU. But for the rest of us it means sulking and scrolling past the announcements, or trying to hunt down a heavily quantized version, or loading it anyway and watching it struggle.