1 article tagged inference.
vLLM from nothing to a working OpenAI-compatible server: WSL2 or Linux, Docker, the compose file we actually run, a preset per model, and the flags that measurably changed things on a dual-4090 bench.