Notes from Running a Local LLM for 3 Weeks

There was a project at work where we couldn't use external APIs, so I ran local models for about three weeks. The conclusion: usable, but not a silver bullet.

I started with a 7B-class model, but the Korean quality was disappointing, so I moved to 13B, and eventually up to the 30B range. The bigger the model, the better the answers, but GPU memory and electricity costs rose along with it. When I ran it quantized, memory dropped sharply, but it felt a bit fuzzier on long contexts.

What I noticed most was inference speed. If only a few tokens per second come out, users drop off quickly. Even if it meant sacrificing a bit of accuracy, prioritizing speed worked better in real use.

Now I split things: simple tasks like summarization and classification run locally, while complex generation goes to the cloud. A hybrid setup seems like the realistic answer.

by 월급루팡334

6 answers

Agreed, hybrid turned out to be the realistic answer. I settled on the same thing.

by 주말개발자677 · ▲0