2 articles tagged benchmarks.
In today's post I'm sharing how I benchmark local models on vLLM by changing one thing at a time, along with the traps that quietly skew the numbers and my results from 21 setups on two 48GB RTX 4090Ds. The harness, vllm-bench, is free, so you can run the same tests on your own card.
local-llm · vllm
In today's post I'm going through what made local LLM inference faster on my two modded 48GB RTX 4090s, in the order it mattered: the file format, how you split a model across cards, locking the clocks at boot, speculative decoding and load times, with the numbers from my own testing.