My Experience Running a Local LLM at Home (7B-Class)

I tried running a quantized 7B model on a 12GB graphics card. To get straight to the point: "It's more usable than I expected, but not quite good enough to be my main setup."

Speed was around 20 tokens per second, which was tolerable, and it did okay with simple summarization or translation. But once you get into reasoning or long-context processing, the difference from API models is clearly big.

Still, the biggest advantage seems to be that it runs without an internet connection and data doesn't leave the machine. I'm considering using it when handling sensitive documents at work.

by 초보개발자554

6 answers

Agreed, the fact that it doesn’t use data is huge. That alone is why I installed it too.

by 지나가던행인114 · ▲0

With 12GB, 7B q4 is just right lol. If you get greedy and go above that, it starts swapping and freezes up.

by 월급루팡142 · ▲0