NoRaincheck

Thinking Local LLMs and AI

Thinking Local LLMs and AI

December 2025

Running models locally is nothing new. Infact I’ve always had a particular affinity to llama.cpp. Recently, there is the newly introduced local text to image (z-image-turbo) generation model that can ‘comfortable’ be run locally (albeit perhaps a bit slow without a dedicated GPU).

Usage would look something like (using justfile to template it) using stable-diffusion.cpp:

 1[no-cd]
 2sd_generate:
 3    PROMPT=$(gum input --placeholder "prompt for image generation"); \
 4    OUTPUT=$(gum input --placeholder "output png file"); \
 5    DYLD_LIBRARY_PATH=/path/to/dyld/library \
 6    sd --difffusion-model z_image_turbo-Q4_0.gguf \
 7    --vae /path/diffusion_pytorch_model.safetensors \
 8    --llm Qwen3-4B-Instruct-2507-Q6_K.gguf \
 9    --cfg-scale 1.0 \
10    --offload-to-cpu \
11    --diffusion-fa \
12    -H 512 -W 512 --steps 9 \
13    -p "$PROMPT" \
14    -o "$OUTPUT"

On M1 Macbook Pro with offload cpu enabled it will take roughly 2 minutes per a step, whereas not offloading will improve performance at the cost of memory consumption (n.b. you should have --offload-to-cpu turned on if you are using a low memory variant).

Similarly, it is possible and perhaps desireable to using llama.cpp via CLI. One random project I want to do with Helix is abstracting over it so that we can use it to generate text by automatically looking up context in directory. This will probably be a mini-project (think of it as an extension of smart-cat)

1$ llama-cli -m /path/to/llm.gguf -n 1 -p "hello how are you" -no-cnv

One of the challenging aspects is if the conversation is multi-turn we would need to format the prompt (-p) to use the chatml template. Having a way to build this up in an intelligent manner is a discussion for another day (is it possible to do it easily via gum with justfile? Maybe?). Another alternative is to make use of aider with helix with the fs_watcher_lsp – this is probably the more natural way to approach it. If we were to implement it ourselves, a way to do this is to imitate the approach taken by aider chat by using Universal Ctags or tree sitter which builds a map for the whole code repository which uses an LLM to figure out which particular file is of interest and then sends the relevant contents to the LLM.

<< Previous Post

|

Next Post >>

🎲 Random post

|

All posts

#LLM #CLI #ML