Local AI agent too slow on Mac mini M4 — how to speed it up
A newcomer running OpenClaw with Ollama on a Mac mini M4 reports painfully slow responses: minutes for web fetches and memory lookups. The likely culprits are a 14B model with a large 16k context window on 24GB of unified memory, plus duplicate Ollama instances competing for the GPU. Practical fixes include shrinking the context window, dropping to a 7B model for the agent loop, and limiting parallel requests.