My Experience Running Local LLMs at Home (7B / 13B / 32B)

I picked up a used 3090 and tried running all sorts of local models. To get straight to the point: if your use case is clear, it’s plenty usable.

Models tested

  • 7B class: Response speed was insane. 60 tokens per second. But reasoning is a bit weak.
  • 13B class: Feels like the best balance. Good enough for summarization and translation.
  • 32B class: Clearly smarter, but 12 tokens per second. If you watch it stream, it’s fairly tolerable.

Pros

1. No internet required. It even runs on a plane.

2. Even if you feed it company code or similar, it never leaves your machine.

3. The psychological comfort of zero token cost is huge.

Cons

1. Electricity bill... leaving it on for a month is only a few thousand won, but in summer the room gets hot.

2. Compared with the latest commercial models, the limitations are clear.

3. In the end, quantization and serving setup take quite a bit of time.

Right now, I use local for summarizing sensitive documents and mix in APIs for everything else. If you have questions, leave a comment.

by 취준생김씨51

3 answers

Is 60 tokens per second from running a 7B at q4? Since the model and quantization level aren't specified, the numbers are a bit ambiguous.

by 문과출신개발자243 · ▲0