NoRaincheck

Local AI is Probably Good Enough

I think local AI on consumer devices is probably good enough now. Price is still an issue, of course, but it’s now possible to run Qwen 3.8 Flash on 64GB of unified RAM (albeit at Q2 or Q3 quants) at a reasonable ~40+ tps with a decent context window. That alone says divorcing yourself from proprietary models is more than possible — and maybe even desirable — even if the models themselves make no further progress.

I do think longer context and better prefill speed are the next ‘big’ items for everyday consumers. Things are still too ‘difficult’, and not ’turn-key’ enough, for the average person.

Definitely looking forward to hardware prices going down.

<< Previous Post

|

Next Post >>

🎲 Random post

|

All posts

#Ai #Local-Llms #Self-Hosting