1 article in Local AI.
I decommissioned Hopper (my local LLM bootstrapped server) and moved my local models to a two-card 4090 rig on vLLM. Much faster - but the tool I wrote to save Claude tokens started handing back blank answers, and it looked like the new rig's fault. It wasn't. Here's the honest version, and how v3.2.1 fixes it.