- · local-llm · gpu
Tuning vLLM on a Modded 48GB RTX 4090: From 18 to 133 Tokens per Second
Three days of benchmarking vLLM on a modded 48GB RTX 4090: the quantisation trap that costs Ada owners 40% of their speed, the free 1.9x from speculative decoding, the KV compression that never served once, and what cutting 120 watts costs you. Every number measured, every failure kept.
Read article - · local-llm · gpu
Best GPUs for Running Local LLMs (2026): Memory Bandwidth, VRAM and the Cards Worth Buying
What to buy in mid-2026 for running local LLMs seriously, from someone who has benched most of it. Why memory bandwidth matters more than FLOPS, why the quantisation format can matter more than the card, and the GPUs worth your money - from the used 3090 floor to the modded 48GB 4090s I ended up buying twice.
Read article