1 article tagged VRAM.
In today's post I'm taking a closer look at the KV cache, the memory a model keeps of the conversation so far, and why 'does it fit in VRAM?' is the wrong question when a local LLM won't load. It's what refused a 70B model its own 131k context on my 48GB card, and I'll show you how to size it, the vLLM settings that control it, and the kilobyte of GPU shared memory that stopped a 16GB model outright.
local-llm · vllm