Review: Running a Local LLM on a Spare GPU at Home
I had a 3060 12GB left over from gaming a while back, so I tried running a local model just for fun. I'd only used paid online services before, but running one myself has its own appeal.
First off, installation wasn't as hard as I expected. I installed Ollama, downloaded one model, and was chatting right away. It probably took less than 30 minutes.
Here's roughly what I thought while using it.
- Privacy is definitely a plus. I can put in internal company documents or personal notes without worrying
- 7B–8B models are a bit awkward at Korean. English is passable, but honorifics often break down
- Speed feels like it's printing one character at a time, so it can be frustrating. Once you go past 14B, it gets noticeably slower
- In the end, I concluded that cloud APIs are better for actual use
Still, it's amazing that we can now do this much offline. Before long, we'll probably be running 20B-class models at home too.
10 answers
For local use, the 3060 12GB has insane value, agreed. I'm also thinking about buying another used one.
The "done in 30 minutes" thing seems like a bit of an exaggeration lol. When you think about having to download tens of gigabytes just for the model at first... Still, it's true that the barrier to entry has dropped a lot.
Starting with Ollama was a good choice. But the awkward Korean from 7B–8B models isn’t because they’re local—it’s just how models at that scale are. If you try EXAONE or Qwen 14B or larger, honorifics are much more stable. With 12GB VRAM, even a 14B model can barely run with 4-bit quantization.