Best PC for Local AI & LLMs 2026: 17 Tested Builds
In today's guide we're working through the best PCs for running local AI and LLMs, from a used RTX 3090 - the card I'd call the sensible minimum - up to Apple's new 512GB M5 Ultra Mac Studio. I've tested 17 builds along the way, so you'll see which hardware runs which model size, and at what tokens per second.
Most people who ask "do I need a powerful PC for AI?" don't need one. If you want to use Claude for daily work - writing, code, analysis, the conversational stuff Claude is great at - the chat client runs on a 2019 MacBook with 8GB of RAM and you'll never notice the machine. Start with the Beginners Guide and come back here when the API bills start mounting or you want to try something the cloud models won't run for you. New to Claude? You can sign up for a free week of Claude Code here .
First, a quick word on what's changed since I last updated this page, because late August 2026 was quite the month. Apple announced the M5 generation of the Mac Studio - the M5 Ultra takes unified memory to 512GB, more than double anything you could put on a desk before, and it ships on 22 September. Street prices on the RTX 5090 have drifted the wrong way again, the 128GB mini PC category has grown a much cheaper entry point, and AMD's next unified-memory chip now has a name and a rough date. All of it is worked into the tiers below.
About two years ago I bought a used RTX 3090. I was just so curious about what an AI model could do, and I happened to have prior experience of GPUs with a bit of gaming PC experience under my belt. I'd pull a 13B model, chat with it, try a bigger one, hit the VRAM wall, and immediately start thinking about a second card. That one card turned into a Threadripper workstation stuffed with Ada GPUs, and eventually into the thing I run now: a pair of modded 48GB RTX 4090s serving through vLLM, with an MCP that hands Claude Code the bulky work so it comes off the token bill. Two years, a lot of wasted money, and the one lesson I'd hand you before you spend a penny: don't bother with anything below a used RTX 3090. The cheaper cards work, technically, but 24GB of VRAM and that card's memory bandwidth are what get you into the models worth running every day - the benchmark tables further down make the case better than I can. What you do need to know is which single number on the spec sheet matters, because most of them don't.
Below: the hardware mistakes and lucky finds, from the used 3090 that started all this to mini PCs with 128GB of unified memory running 70B models from under a desk. Every number I could measure on my own kit, I have; where a card isn't mine, the benchmark is named and credited - which, as far as I can tell, is more than any other page on this subject bothers to do.
Quick Navigation
Jump directly to what you're looking for:
Skip to the PC Finder | VRAM & Quantisation | Hardware Secrets | Budget (Under £800) | RTX 3060 Benchmarks | Mini PCs for AI | Mid-Range (£800-£3,000) | Premium (£3,000+) | Does the CPU Matter? | Software Stack | Comparison Table
Just want a recommendation? Use the PC Finder
If you already know roughly what you want to do (run a 70B model, generate images, drive a local code agent) and how much you want to spend, scroll the picker below. It filters 17 hand-checked builds from Corsair, Apple, Framework and NVIDIA across six tiers. Prices come from live merchant feeds and we re-check them every few weeks.
If you want to understand why a particular tier suits you, keep reading after the finder. The article is structured so you can stop at the recommendation that fits, or keep going for the full hardware reasoning.
Find an AI PC that fits your budget and your models
Filter by what you want to run, what you want to spend, and the form factor you have room for. Showing 17 of 17 hand-picked builds.
Filters
Disclosure Prices verified against merchant feeds on the date above. They drift weekly - check before you buy. Affiliate links earn Houtini a small commission at no cost to you and fund the next round of testing. Last verified 2026-09-01. Dual-currency display uses an approximate GBP/USD rate of ~1.27 (2026-09-01); real rate, shipping, and import duties not included.
-
In stock
CorsairCorsair VENGEANCE i7600 (RTX 5070)
The Corsair entry point - 12GB VRAM covers the 7B-13B models
Price (US listing)$2,400~£1,890GPU NVIDIA RTX 5070 VRAM 12GB CPU Intel Core Ultra 7 265KF RAM 32GB SSD 1TB12GB of VRAM runs Llama 3.1 8B at Q4 with room for context, and SDXL for image work. 13B fits if you drop the quantisation. The ceiling arrives before 30B, so if bigger models are the plan, start a tier up instead - this is the box for finding out whether local AI earns a place in your week.
via Corsair View at Corsair -
In stock
CorsairCorsair VENGEANCE a5100 (RTX 5070 Ti)
16GB VRAM and the 9800X3D - the sensible middle of the range
Price (US listing)$3,200~£2,520GPU NVIDIA RTX 5070 Ti VRAM 16GB CPU AMD Ryzen 7 9800X3D RAM 32GB SSD 2TB16GB VRAM is where the mid-size models come into reach: Mistral Small 24B at Q4 with a 16K context, Qwen coder models at Q3. The 9800X3D chews through prompt processing, which you notice when you paste a few thousand tokens of code at a local model. 2TB of storage means the model library has somewhere to live.
via Corsair View at Corsair -
In stock
CorsairCorsair VENGEANCE a7500 AIR (RTX 5080)
Newest X3D silicon with the RTX 5080 - fast 16GB with headroom
Price (US listing)$3,600~£2,835GPU NVIDIA RTX 5080 VRAM 16GB CPU AMD Ryzen 7 9850X3D RAM 32GB SSD 2TBSame 16GB VRAM ceiling as the 5070 Ti build, but the 5080's extra compute shows on image generation and long-context prompt processing, and the 9850X3D is the current top single-CCD X3D part. The AIR chassis trades liquid cooling for simplicity - a little louder under sustained load, one less thing to maintain.
via Corsair View at Corsair -
In stock
CorsairCORSAIR ONE a600 (RTX 5080)
Liquid-cooled compact tower - the quiet route to a 5080
Price (US listing)$4,800~£3,780GPU NVIDIA RTX 5080 VRAM 16GB CPU AMD Ryzen 7 9800X3D RAM 32GB SSD 2TBLiquid-cooled CPU and GPU in a compact tower that sits on the desk rather than under it. You pay a premium over the VENGEANCE builds for the same silicon; what you get back is near-silence under sustained inference - which matters a great deal if the machine shares a room with you all day.
via Corsair View at Corsair -
In stock
CorsairCorsair VENGEANCE i7500 (RTX 5090)
The cheapest Corsair route to 32GB of VRAM
Price (US listing)$5,700~£4,488GPU NVIDIA RTX 5090 VRAM 32GB CPU Intel Core i9-14900KF RAM 64GB SSD 2TB32GB of VRAM is the number that changes what you can run: Llama 3.3 70B fits at Q4 entirely on the GPU, with 64GB of system RAM behind it for hybrid experiments. The i9-14900KF is a generation older than Corsair's X3D builds and slower at prompt processing, but for GPU-bound inference - which most local LLM work is - the 5090 is doing the lifting.
via Corsair View at Corsair -
In stock
CorsairCorsair MILLENNIUM (RTX 5090)
Top X3D silicon with the 5090 and 4TB of storage
Price (US listing)$7,000~£5,512GPU NVIDIA RTX 5090 VRAM 32GB CPU AMD Ryzen 9 9950X3D RAM 64GB SSD 4TBThe 9950X3D is the top dual-CCD X3D part, and paired with the 5090's 32GB it covers 70B Q4 inference with the CPU cores to run parallel agents alongside. 4TB of NVMe holds a real model library. If the 192GB-RAM a7500 AIR flagship is more than you need, this is the same class of machine at a saner price.
via Corsair View at Corsair -
In stock
CorsairCorsair VENGEANCE a7500 AIR
Prebuilt flagship - RTX 5090 + 192GB DDR5 + 9950X3D in a Corsair chassis
Price (US listing)$7,000~£5,512GPU NVIDIA RTX 5090 VRAM 32GB CPU AMD Ryzen 9 9950X3D RAM 192GB SSD 6TB192GB of DDR5 system RAM is the killer feature here - that's enough headroom to run hybrid CPU+GPU inference on 200B+ models that won't fit on a single 5090 alone. Top-spec 9950X3D handles the prompt processing; 6TB of NVMe across three drives means you keep a real model library. Premium build quality, full Corsair component stack, and the price reflects all of it.
via Corsair View at Corsair -
In stock
Empowered PCEmpowered PC Panorama XL
US Amazon-stocked flagship - 9950X3D + RTX 5090 + 128GB DDR5 in one box
Price (US listing)$7,699~£6,062GPU NVIDIA RTX 5090 VRAM 32GB CPU AMD Ryzen 9 9950X3D RAM 128GB SSD 10TBSticker is $7,699 on Amazon US. The GBP shown is a straight FX conversion - UK buyers add ~£1,200 for typical Amazon import duty + shipping on a desktop tower this size. 16-core X3D plus 32GB VRAM 5090 plus 128GB DDR5 is the most balanced flagship workstation I've seen on Amazon - the 128GB system RAM means CPU offload for 100B+ models moves from possible to pleasant.
via Amazon View at Amazon -
In stock
AppleApple Mac Mini M4 (16GB)
The cheapest way to dabble with on-device AI - not for serious LLM work
Price£599~$761GPU Apple M4 GPU (10-core) VRAM 16GB CPU Apple M4 (10-core) RAM 16GB SSD 0.3TBApple's unified memory means the 16GB is GPU-addressable. Llama 3.1 8B at Q4 runs at ~25 tokens/sec. Not enough for 30B, no headroom for context-heavy work. Tiny, silent, low power.
via Apple View at Apple -
In stock
AppleApple Mac Mini M4 Pro (64GB)
Punches above its size - 30B models in a box smaller than a hardback book
Price£2,199~$2,793GPU Apple M4 Pro GPU (20-core) VRAM 64GB CPU Apple M4 Pro (14-core) RAM 64GB SSD 1TB64GB unified means Llama 3.3 70B runs at Q4 with 8K context, around 8 - 10 tokens/sec on MLX. Slower than a 5090 but in a fraction of the space at a fraction of the power draw. Best mini-PC option for local LLMs full stop.
via Apple View at Apple -
In stock
AppleApple Mac Studio M4 Max (64GB)
Desktop Apple silicon with the GPU cores to match the unified memory
Price£3,499~$4,444GPU Apple M4 Max GPU (40-core) VRAM 64GB CPU Apple M4 Max (16-core) RAM 64GB SSD 1TB40 GPU cores means 70B Q4 runs noticeably faster than on the Mac Mini Pro - closer to 14 tokens/sec. The Studio is the right Apple form factor if you want sustained AI workloads without thermal throttling.
via Apple View at Apple -
In stock
AppleApple Mac Studio M4 Max (128GB)
128GB unified fits the big models entirely - 100B+ at Q4
Price£4,999~$6,349GPU Apple M4 Max GPU (40-core) VRAM 128GB CPU Apple M4 Max (16-core) RAM 128GB SSD 1TBLlama 3.1 405B fits at Q2 with 4K context. Qwen 2.5 72B at Q8. Mistral Large 123B at Q4. This is where Apple's unified memory model beats a 5090 - you trade speed for fitting models that just won't load on 32GB VRAM.
via Apple View at Apple -
In stock
AppleApple Mac Studio M3 Ultra (256GB)
The biggest unified memory box money buys - frontier-scale models locally
Price£8,499~$10,794GPU Apple M3 Ultra GPU (80-core) VRAM 256GB CPU Apple M3 Ultra (32-core) RAM 256GB SSD 1TBLlama 3.1 405B at Q4. DeepSeek V3 671B at Q3. The only consumer box that runs frontier-class models locally. Slower per-token than a 5090 but the 5090 can't load these at all. If you're trying to run what OpenAI and Anthropic ship, this is the desk-side option. One thing before you order: Apple announced the M5 Ultra on 25 August (ships 22 September, up to 512GB unified) - if you can wait a few weeks, wait, and watch M3 Ultra clearance prices while you're at it.
via Apple View at Apple -
In stock
CorsairCorsair AI Workstation 300
Strix Halo in a 4.4-litre Corsair chassis - 70B in a quiet box
Price (US listing)$2,700~£2,126GPU Integrated Radeon 8060S (40 CU) VRAM 96GB CPU AMD Ryzen AI Max+ 395 RAM 128GB SSD 1TBSame Strix Halo silicon as the Framework Desktop in Corsair's own 4.4-litre chassis with a 300W Flex ATX PSU. 96GB of the 128GB unified pool is allocatable to the integrated GPU - enough to run Llama 3.3 70B at Q4 with comfortable context. Less modular than the Framework but the Corsair cooling is properly tuned for sustained inference loads.
via Corsair View at Corsair -
In stock
GMKtecGMKtec EVO-X2
First Strix Halo mini PC to market - Amazon UK direct, 128GB unified
Price£2,099~$2,666GPU Integrated Radeon 8060S (40 CU) VRAM 96GB CPU AMD Ryzen AI Max+ 395 RAM 128GB SSD 2TBThe surprise entry that opened up the Strix Halo mini PC category. Same chip and unified memory as the Framework and Corsair builds - gets to market via Amazon UK direct, which makes it the easiest path for UK buyers who don't want to deal with Framework's order queue or Corsair's US-to-UK shipping. 2TB NVMe at the entry tier vs 1TB on the Corsair is a small but real edge.
via Amazon View at Amazon -
In stock
FrameworkFramework Desktop (Ryzen AI Max+ 395, 128GB)
Mini-ITX with 128GB unified-style memory - the AMD answer to a Mac Studio
Price£1,970~$2,502GPU Integrated Radeon 8060S (40 CU) VRAM 128GB CPU AMD Ryzen AI Max+ 395 (16-core) RAM 128GB SSD 1TBAMD's Strix Halo silicon shares system RAM with the integrated GPU - same trick Apple uses. 128GB is allocatable to the GPU. Inference speed is lower than discrete GPUs (~6 - 8 tokens/sec on 70B Q4) but you fit models that won't run on a 5090. LPDDR5X supply pushed the price up from the $1,999 launch, but it still undercuts a Mac Studio 128GB.
via Framework View at Framework -
In stock
NVIDIANVIDIA DGX Spark
NVIDIA's first desktop AI workstation - 128GB unified GPU memory
Price£3,799~$4,825GPU NVIDIA Blackwell GPU VRAM 128GB CPU NVIDIA GB10 Grace ARM (20-core) RAM 128GB SSD 4TBNVIDIA's desktop entry to the full CUDA stack: 128GB unified between the Grace ARM CPU and Blackwell GPU, running 200B-class models at Q4. Launched at $3,999 and hiked to $4,699 in February 2026 on memory supply constraints. ARM Linux only - it won't run Windows apps, so it's a dedicated AI box rather than a daily driver.
via NVIDIA View at NVIDIA
Nothing matches every filter. Loosen one - most people drop the budget filter first.
The picks above are deliberately narrow. There are hundreds of PCs that could technically run a local model. These are the ones I’d buy or recommend to a friend at each tier, with the trade-offs called out in the notes on each card. The rest of this article explains why VRAM is the single number that matters and where each tier hits its ceiling.
The number that matters
Something I’ve learned building these rigs: on a PC workstation, VRAM on the GPU decides everything. Not clock speed, not CUDA cores, not the number NVIDIA puts on the box. If a model fits in your GPU’s memory, it runs fast enough. If it doesn’t fit, you’re at two tokens per second because your machine will most likely attempt to fit the rest of the model in your main RAM - it still works but it’s slooow!
People overcomplicate this. Here’s the bare maths: every parameter in your model has to sit in memory somewhere - you can’t get around that. At full precision (FP16), one parameter costs you 2 bytes. A 70 billion parameter model at full precision is 140GB. No consumer GPU on the planet has that kind of VRAM, yet. Even 3 of those 48GB modified “4090D” cards you see on eBay would probably melt. (There are “4090 D’s” on eBay that have been reboarded to accommodate 48GB. I was so, so tempted that I eventually bought two, at £3,100 each - they’re the pair running this site’s benchmarks through vLLM right now. The boards come out of the same factories as the NVIDIA cards, they swap over the GPU chip and add better RAM. A lot less sketchy than you’d think.)
Quantisation fixes this. Compress those billions of parameters down to 4-bit (Q4_K_M is the format you’ll see everywhere) and each one drops to roughly half a byte. That 70B model goes from 140GB to about 40GB - two used RTX 3090s with room for context window overhead. I’ve been running Qwen 3 Coder Next at Q6 quantisation on my own rig for about six months now and can’t feel any quality difference from full precision on the tasks I throw at it. I wrote up the whole process in my LM Studio setup guide if you want to try the same thing on your own rig.
One thing that’s changed since I first wrote this article: Llama 4 Scout landed with 109B total parameters in a Mixture-of-Experts architecture. Only 17B parameters are active at any time, but MoE models need all parameters loaded in memory. That means ~55-70GB of VRAM just to load it at INT4. A single RTX 4090 can’t touch it. This is pushing people toward either high-RAM unified memory machines (the mini PCs below) or multi-GPU rigs. The VRAM arms race isn’t slowing down.
Quick VRAM guide
So what sort of size model can run on your VRAM? Beware - 3080’s have 10GB versions so watch out if you’re buying second hand on eBay.
| Model Size | At Q4 (4-bit) | At Q8 (8-bit) | Good GPU Fit |
|---|---|---|---|
| 7B | 6-8 GB | 10-12 GB | RTX 3060 12GB, RX 9060 XT 16GB |
| 13-14B | 10-12 GB | 16-18 GB | RTX 3060 12GB, RX 9060 XT 16GB, 4060 Ti 16GB |
| 34B | 20-24 GB | 30+ GB | RTX 3090 or 4090 (24GB) |
| 70B | ~40 GB | ~75 GB | Dual 3090s, Mac Studio, RTX 5090, or 128GB mini PC |
| 109B MoE (Llama 4 Scout) | ~55-70 GB | ~110+ GB | 128GB unified memory (Framework/GMKtec/DGX Spark) |
| 100B+ dense | 60-70 GB | 100+ GB | Quad 3090s, M3 Ultra 192GB |
Don’t forget KV cache on top of this - it stores your conversation state and grows with context length. At 32k tokens, budget for another 2-4GB, which caught me out the first time I tried to squeeze a 34B onto a 24GB card.
Local LLM hardware secrets
Here are three bits of hardware wisdom I’ve had to learn the expensive way.
One fast card beats two slow ones
Tempting maths: two RTX 3060 12GB cards give you 24GB total. Same VRAM as a single 3090. Same capacity, completely different speed. This is a big mistake on my part - I bought an array of bargain Ada generation RTX 4000s and 4500s - the mixture of the cards and the volume of them was a mistake. It runs but I think I’m losing at least 20% of the performance because of all the PCI lanes in play.
Digital Spaceport did the numbers. A single 3090 hits 28 tokens per second on Gemma 3 27B at Q4. The dual 3060 setup? Six. On the exact same model. Splitting a model across GPUs over PCIe - sharding, they call it - kills throughput because the cards spend more time talking to each other than doing inference. I really, really wish I understood this before I sold my 3090’s from my GPU mining days.
So when does multi-GPU work? When the model needs both cards anyway. Two 3090s running a 70B model that requires 48GB of VRAM is fine - fifteen to twenty tok/s with NVLink. But don’t buy two cheap cards hoping they’ll match one expensive one. They won’t. Don’t mix generations of cards, don’t mix VRAM numbers - and most consumer “gaming” motherboards don’t support full 16-channel PCI on more than one of the PCI slots. Simple is the best approach.
NVLink for Multi-GPU fine tuning
Something I didn’t expect when I added my second GPU: the connection between the cards ends up mattering almost as much as the cards themselves in specific use cases. What NVLink does is give the GPUs their own private highway - 112.5 GB/s bidirectional. Compare that with regular PCIe 4.0 x8, which tops out around 16 GB/s. About seven times slower, and you notice it in practice.
The caveat: NVLink is better for fine tuning performance, not for inference (chat!) - oh well.
Apple Silicon: capacity over speed
A Mac Studio M3 Ultra with 192GB of unified memory can load models that would need four discrete NVIDIA GPUs on a PC - and the newly announced M5 Ultra pushes that ceiling to 512GB (more on it in the premium tier). All that RAM is GPU-accessible. No PCIe bottleneck, no sharding penalty. Near-silent, too, which matters if (like me) you’re working in the same room as the hardware.
Speed-wise, NVIDIA is quicker - about 2-3x on models that fit in its VRAM. A dual 3090 PC does 15-20 tok/s on 70B; the M3 Ultra manages 8-12 tok/s on the same model. Where the Mac pulls ahead is models above 100B parameters that the PC can’t touch without a quad-GPU build, and frankly, for research tasks where you’re running huge models rather than chatting interactively, it makes more sense than people give it credit for.
Unified memory changes everything
So here’s what changed in 2026. Apple proved the concept with Apple Silicon years ago, but now AMD’s Strix Halo chips bring the same unified memory architecture to Windows PCs. The Framework Desktop, Corsair AI Workstation 300, and GMKtec EVO-X2 all pack 128GB of shared memory that the integrated GPU can access directly. No PCIe bus, no sharding. You load a 70B model into memory and the GPU just… uses it.
The trade-off is speed. These integrated GPUs are slower than a dedicated RTX card on models that fit in discrete VRAM. But for models that don’t fit - 70B, Llama 4 Scout, anything MoE - unified memory machines are the only option under £3,000 that doesn’t involve multiple GPUs and a wiring diagram. I got into all of this in my beginner’s guide to AI mini PCs and the DGX Spark - worth reading if the unified memory thing is new to you.
Budget: Under £800
A line in the sand first, because it will save you money: the used RTX 3090 is where I’d start, and honestly, I wouldn’t bother with anything below it. The cheaper cards further down run small models perfectly well - there’s a benchmark table proving it - but 24GB of VRAM and 936 GB/s of memory bandwidth is the point where local AI stops being a novelty, and the cards below that line are the kind of purchase you end up making twice.
RTX 3090 24GB (Used/Renewed)

Yeah, it’s two generations old. The local AI community collectively shrugged at that ages ago and kept buying them.
Twenty-four gigabytes of GDDR6X handles 34B models at Q4 or 70B at tight quantisation. I ran one of these for about a year before the workstation build happened, and looking back I'm slightly embarrassed at how long I underestimated what a single 24GB card could handle. Community benchmarks from Digital Spaceport show 28-36 tok/s on 14B models, 28 tok/s on Gemma 3 27B Q4. Nothing under a grand comes close to that combination of capacity and speed, and it's got NVLink support for when you inevitably want to add a second one. Used prices have settled around £380-480 in mid-2026, which makes a dual-3090 48GB setup achievable for under a thousand pounds - still the value king of local AI.
Renewed cards run £650-800 on Amazon . Most sellers give you about 90 days of warranty. Bit of a gamble, but I’ve not heard of widespread failure rates from the AI community. If you’re planning multi-GPU later, look for blower-style cards - they exhaust heat out the back instead of dumping it onto the card above. The 350W TDP per card adds up fast when you’ve got two of them in the same case.
NVIDIA RTX 3090 24GB (used)
- VRAM 24GB GDDR6X · 936 GB/s
- Best for up to 34B at Q4; 70B with a second card
What about the RTX 3060 and AMD’s RX 9060 XT?

The RTX 3060 12GB used to open this guide as the £200 way in, and to be fair, it still does what it always did: 7B models at Q8, 13-14B at Q4, all the capable smaller models running at usable speeds. The reason it opened the guide is that 12GB from an older generation beats the 8GB on a newer RTX 4060 for AI work, every time - that part hasn’t changed.

AMD’s RX 9060 XT 16GB is the newer budget option at around £300. Four extra gig over the 3060 means 14B models at Q8 fit comfortably, and ROCm support has improved hard over the past year - Ollama, llama.cpp and LM Studio all run on AMD now - but you’ll still hit more edge cases than you would on NVIDIA.
So why don’t I recommend either of them any more? Because of what happens the first time you want a model that doesn’t fit. Both cards run out of road at about 20B parameters, and past that point the model spills into system RAM and your tokens per second collapse to single digits. The benchmarks below show exactly where that wall is - and the 3090’s 24GB, with nearly a terabyte per second of memory bandwidth behind it, is what moves the wall far enough away that you stop thinking about it.
RTX 3060 12GB benchmarks
People keep searching for specific numbers on this card, so here they are - community benchmarks from Hardware Corner and Digital Spaceport:
| Model | Quantisation | Tokens/sec | Context |
|---|---|---|---|
| Llama 3 8B | Q4_K_M | ~42-50 | 4k-16k |
| Mistral 7B | Q4_K_M | ~40-50 | 4k-16k |
| Qwen 2.5 14B | Q4_K_M | ~22-23 | 16k |
| Qwen 2.5 14B | 5-bit EXL2 | ~30-33 | 8k |
| Phi-3 14B | Q4_K_M | ~22-25 | 16k |
| Any 20B+ model | Q4 | ~9 | Limited |
The 14B sweet spot is the surprise. Thirty tokens per second on Qwen 2.5 14B at 5-bit EXL2 is quick enough to forget you’re on local hardware. But look at that last row, because it’s the whole argument for spending more: past 20B parameters the model spills into system RAM and drops to single digits. Still usable if you’re batching things overnight, but not for a conversation - and the models worth talking to keep getting bigger.
If you’re running ExLlamaV2 (which you should be for GPU-only inference on NVIDIA), the 360 GB/s memory bandwidth on the 3060 outperforms the RTX 4060 on token generation. Newer architecture doesn’t matter when you don’t have enough VRAM.
Mini PCs for AI
This is the section that didn’t exist when I first wrote this article, and it’s the biggest single change since the first version. Everything moved when AMD shipped Strix Halo - a laptop-class chip with 128GB of unified memory that the integrated GPU can access directly. Suddenly you can run 70B models from a box that fits on a shelf and draws 120W. No discrete GPU needed.
I covered the technology in depth in my beginner’s guide to AI mini PCs and the DGX Spark , but here’s the practical buying guide.
Framework Desktop (128GB)

If I were starting from scratch today this is probably where my money would go. AMD Ryzen AI Max+ 395, 128GB LPDDR5X unified memory, crammed into a 4.5-litre case that Framework co-designed with Cooler Master and Noctua. And because it’s Framework, the whole thing is modular - you can swap the front panel tiles, the fans, even 3D print custom bits.
96GB of that 128GB is allocatable to the GPU. Runs 70B models. Llama 4 Scout fits (just). Near-silent under inference load.
The catch: LPDDR5X prices have gone through the roof. Framework originally priced the 128GB model at $1,999 but it’s now up to around $2,459 (~£1,970) due to memory supply constraints. Still the cheapest 128GB unified memory machine you can buy, and the modular design means you’re not throwing the whole thing away when the next generation of chip arrives.
Corsair AI Workstation 300

Corsair’s answer to the same question, and it comes in tiers rather than one box. The one you can buy today is the $1,699 model: a Ryzen AI Max 385, 64GB of LPDDR5X, and up to 48GB of that addressable as VRAM, all in Corsair’s own 4.4-litre chassis with a 300W Flex ATX PSU. The one you’ll really want is the $2,699 step up - the Ryzen AI Max+ 395 with the full 128GB and up to 96GB usable as VRAM, which is where 70B models start to fit without a fight - but as I write this it’s out of stock, so the card below is the in-stock entry point.
Less modular than the Framework, and you’re buying into Corsair’s ecosystem rather than a repairable box. What you get for that is proven build quality and cooling, plus the bundled Corsair AI Software Suite, which does more of the setup hand-holding than the Framework’s bare-metal approach. If you already trust Corsair hardware, it’s the path of least resistance into unified memory.
- Chip AMD Ryzen AI Max 385 (8C / 16T)
- Unified memory 64GB LPDDR5X-8000
- Usable as VRAM up to 48GB
- NPU XDNA 2 · 50 TOPS
- Chassis 4.4L · 350W · USB4 · WiFi 6E
GMKtec EVO-X2

The surprise entry that started this whole mini PC category. Same AMD Ryzen AI Max+ 395, same 128GB option, slightly different cooling approach. Around £2,000-2,500 on Amazon .
It was the first to market and the early reviews are solid. Speed won’t match a 3090, mind you - somewhere around 10-15 tok/s on 27B models from what I’ve seen. For an always-on inference box that handles 70B from under your desk without waking the house though, I haven’t found anything else in this bracket. Plus I expect to see gen 1 Strix Halo mini PCs on eBay for £500-600 in two years’ time once the next chip generation lands.
BOSGAME M5
The cheapest way into the 128GB club as I write this. The BOSGAME M5 carries the same Ryzen AI Max+ 395 as the Framework and the GMKtec, with the full 128GB of LPDDR5X, and it has been dipping under $1,700 in recent sales - several hundred pounds below the Framework for the same silicon. BOSGAME is a smaller brand and the support story reflects that, so weigh the saving against the warranty question, but Wccftech flagged it as the cheapest 128GB Strix Halo box going and nothing I’ve seen since has undercut it.
One more thing to watch out for before you commit to any of these boxes: AMD’s follow-up chip, Medusa Halo, is now on the roadmap for 2027 - Zen 6 cores, RDNA 5 graphics and LPDDR6, which should lift the memory bandwidth this whole first generation is limited by. If you need the machine now, buy it now; if you’re merely curious, the second generation is the one that fixes the first generation’s main complaint.
Framework Desktop
- 128GB unified memory
- AMD Strix Halo
- 4.5L case
“Modular, repairable”
Corsair AI Workstation 300
- 64-128GB unified memory
- AMD Strix Halo
- 4.4L case
“64GB config in stock today”
GMKtec EVO-X2
- 128GB unified memory
- AMD Strix Halo
- Mini PC form factor
“96GB allocatable VRAM”
BOSGAME M5
- 128GB unified memory
- AMD Strix Halo
- Mini PC form factor
“Cheapest 128GB entry”
NVIDIA DGX Spark
- 128GB unified memory
- Grace Blackwell
- Desktop form factor
“1 petaFLOP FP4”
Mid-Range: £800 - £3,000
Mac Mini M4 Pro (24GB)

Apple’s cheapest route into unified memory for AI work. Twenty-four gig of unified memory, which handles 13-14B models nicely through MLX. Slower on raw tok/s than a 3090, but the software side is painless - Ollama runs natively, no CUDA drivers to wrestle with. £1,399 on Amazon .
Not going to touch 70B, not remotely. But for 7-14B work - coding assistants, summarisation, local chatbots - a lovely quiet machine that does exactly what you’d want. If you’re on macOS already and want to dip a toe into local inference, this is probably where I’d point you first. One thing to check before you order, though: Apple announced the refreshed Mac mini with the M5 Pro in late August 2026, from $1,699 - so make sure of which generation you’re being sold, and expect the M4 Pro to get cheaper as the new one lands.
RTX 4000 Ada (Workstation, 20GB)

I ran these for a good while and they were brilliant for a machine that sits next to you all day. Single-slot form factor at 130W per card, twenty gig of VRAM each. Stick four of them in a standard workstation case and you're sitting on 80GB total at 520W combined, more than enough for 70B models at Q5 with headroom left over for context windows.
I ran six of them in a Threadripper workstation (mixed with RTX 4500 Adas) for 104GB total, and the thing I cared about most was that it was quiet enough to sit in my office all day while I worked next to it. The whole system pulled about 800W under full inference load, which sounds like a lot until you compare it with a quad-3090 setup drawing 1,400W. Raw tok/s per card is lower than gaming GPUs, but the density and near-silence are what sold me for a machine that ran continuously. I have since consolidated to a pair of modded 48GB 4090s on vLLM, but for someone who wants dead-quiet density rather than raw speed, this is still the route I'd point them at. About £1,150 each on Amazon .
Dual RTX 3090 build
The prosumer sweet spot for people who want 70B models on NVIDIA hardware. Two 3090s together give you 48GB of total VRAM. Bridge them with NVLink and you’re looking at 15-20 tok/s on 70B Q4. Skip the bridge and it drops to 10-14 tok/s, which sounds bad until you try it - still plenty fast enough to hold a conversation with a model.
Build essentials: the pair of cards will set you back £1,300-1,500 used. The PSU situation gets interesting because each card wants 350W under load, so budget for a 1,200-1,600W unit. For the platform, Threadripper or HEDT gives you full x16/x16 PCIe bandwidth - consumer boards like Z790 or X670E split to x8/x8, which works but costs some throughput. An NVLink bridge runs about £40-60 used. Whole thing comes in at £1,800-2,200 depending on your platform choice, and no, nobody sells this as a pre-built - you’re getting your hands dirty.
Premium: £3,000+
RTX 5090 (32GB)

The biggest single card you can walk into a shop and buy. Thirty-two gigabytes of GDDR7, 512-bit bus, Blackwell architecture - and for the first time, a quantised 70B model fits on one card. No sharding, no NVLink, no dual-GPU headaches. One slot, done.
Bad news on pricing, though, and it has got worse rather than better. The £1,799 MSRP is a fantasy at this point - GDDR7 supply constraints and AI demand have pushed US street prices to $4,000-4,800 as of September 2026, with UK cards from around £3,240 for the cheapest models up to £3,500-3,800 for the premium ASUS, MSI and Gigabyte boards. Used cards hover around £2,700 on eBay. Budget £3,250 minimum, check the price on the day, and don’t expect it to improve in the immediate term.
The 575W TDP is substantial, too - make sure your PSU can handle it before you get excited and order one.
Interesting side note: Gigabyte launched the AORUS RTX 5090 AI Box - an external GPU enclosure with Thunderbolt 5 that’s specifically marketed for AI workloads. If you’ve got a laptop with Thunderbolt 5, you could run 70B models through an external box. Haven’t tested it myself, but for a laptop-tethered AI workstation the architecture makes sense - Thunderbolt 5 at 80 Gb/s is enough bandwidth to keep a single-card inference workload fed.
NVIDIA DGX Spark

NVIDIA’s “personal AI supercomputer” that I covered in detail in my beginner’s guide to AI mini PCs . The Grace Blackwell GB10 chip with 128GB of unified LPDDR5X and up to 1 petaFLOP of FP4 performance. This is the premium version of the same unified memory concept as the Strix Halo mini PCs above, but with NVIDIA’s own silicon and full CUDA stack.
Originally launched at $3,999, but NVIDIA hiked the price to $4,699 (~£3,800) in February 2026 due to memory supply constraints. Available on Amazon and direct from NVIDIA. If you want 128GB of unified memory with NVIDIA’s ecosystem rather than AMD’s, this is it - but you’re paying a significant premium over the Framework Desktop for that CUDA compatibility.
RTX 4090 (24GB)

Still the fastest card with 24GB of VRAM, and by a decent margin over the 3090 on raw tok/s. Same VRAM ceiling though, and that’s the catch - twenty-four gig is twenty-four gig regardless of what you paid. Buying new? Get this one. Buying used? The 3090 at roughly half the price gives you the same model capacity - which is the metric that matters for local AI. £1,600-2,000 on Amazon .
Mac Studio: the M5 generation

For running the biggest models money can buy in a desktop form factor. Last time round I said that if you were considering an Ultra and could wait, you should, because the M3 Ultra would drop in price the moment its successor shipped. That moment has arrived: Apple announced the M5 generation on 25 August 2026, and the Mac Studio M5 Ultra ships on 22 September with up to 512GB of unified memory and 1.2TB/s of memory bandwidth. Half a terabyte of GPU-accessible memory in a quiet box on your desk, when the previous desktop ceiling anywhere was 192GB.
The M5 Max sits below it with up to 128GB at 614GB/s, which covers 70B models with room for long context. And the M3 Ultra at 192GB is now exactly the discount play I hoped it would be - watch refurb and clearance pricing as M5 stock lands, because 192GB of falling-price unified memory is suddenly the value story at the top of this table.
What the 512GB ceiling means in practice: the very large MoE models - the 200B-400B class that until now belonged in server racks - become loadable on a machine you order from apple.com. Slower per token than NVIDIA on anything that fits in NVIDIA VRAM, same as ever, but there is simply no other way to get this much model into one quiet box.
Quad RTX 3090 build (AM4/AM5)
Digital Spaceport validated this build: four RTX 3090s on an AM4 B550 motherboard. Ninety-six gigabytes of VRAM. We’re talking 100-180 tok/s on 12-20B models, which is absurd throughput. Price per GB of VRAM works out to roughly £30/GB - the cheapest path to serious capacity if you don’t mind some noise.
The PSU needs to be a 2,000W unit minimum and you’ll want a case with serious airflow (or an open-air test bench, which is what most people building these seem to end up with). Fair warning: your partner will comment on the noise. You’ve basically built a small datacenter that happens to live under your desk. Budget: £3,000-3,500 for GPUs plus platform.
Does the CPU matter?
Short answer: not much for inference, and I say this as someone who has run everything from a consumer board to a Threadripper workstation for exactly this. The GPU does almost all the work during token generation. Where the CPU matters is prompt processing (the initial “thinking” phase before the model starts responding) and if you’re offloading layers to system RAM because your model doesn’t quite fit in VRAM.
For a dedicated inference machine, any modern 6-core CPU is fine. Don’t spend £500 on a CPU when that money could go toward more VRAM. The one exception is the unified memory machines (Framework Desktop, DGX Spark) where the CPU and GPU share memory bandwidth - there, the chip choice is the whole machine.
One thing the spec sheets bury, and it only bites once you go past a single card: a gaming motherboard gives you one true PCIe x16 slot. Drop a second GPU in and the board splits the lanes, so the second slot runs at x8, sometimes x4. The counterintuitive part is what it costs you - it slows model loading, not inference. Once the weights are sitting in VRAM the GPU works from its own memory and the link width barely registers; it's only the initial copy from system RAM that feels a narrow slot. A handful of consumer boards, short of full server territory, will bifurcate the main slot to x8/x8 and keep both cards fed evenly - and that, give or take, is the practical line between a standard AM5 build and stepping up to a Threadripper platform with the lanes to spare.
Software stack
Buying the hardware is the easy bit - it’s the software stack where people tend to get stuck. The tools below are a ladder, and you climb it as your ambitions grow.
Start here: Ollama and LM Studio
Ollama - installed it the day I got my first 3090, and it’s still the one I’d tell anyone to start with. The whole workflow is ollama pull llama3:70b and then you’re chatting. Quantisation handled for you, works on everything. Benchmarks I’ve looked at suggest you lose maybe 10-30% on raw throughput versus running llama.cpp bare - which sounds bad until you remember Ollama had you running models in five minutes flat while you’d still be reading llama.cpp compile flags.
LM Studio has the best GUI experience I’ve found for local models. Built-in model browser, chat-with-your-files (that’s RAG), no terminal needed. Perfect if terminals make you nervous. LM Studio ran my inference rig before the move to vLLM, and it still pairs with the houtini-lm MCP server I built for offloading work from Claude Code to cheaper models. I also wrote a full setup guide if you want to get started.
The power tools: llama.cpp and text-generation-webui
llama.cpp is the speed baseline that everything else gets measured against. More config, more control, faster output. Serious multi-GPU setups tend to run this directly rather than going through Ollama’s wrapper.
text-generation-webui - oobabooga’s project, and the one that taught me most about how inference works. You pick between ExLlamaV2 (fastest GPU-only loader) or llama.cpp (flexible CPU offloading) depending on your hardware situation. Learning curve is real, took me a solid weekend to get comfortable, but once you’re past that you can tune everything and understand why your settings matter.
The top rung: vLLM
vLLM is the rung above all of these, and it’s where my own rig ended up. It’s a production inference server rather than a desktop app - continuous batching, paged attention, proper concurrency - and on the same hardware it can be dramatically quicker than the desktop tools once you’re serving real workloads. The setup is Docker and config files rather than a friendly GUI, and the gotchas are real (quantisation-format support on consumer cards bit me more than once). I’ve written the whole journey up: how to set up vLLM , the 48GB 4090 tuning guide , and the twelve-model benchmark runbook with every number measured on the pair of modded cards this article keeps mentioning. Start with Ollama or LM Studio; graduate to vLLM when you catch yourself caring about tokens per second.
GGUF vs EXL2
Two model formats you’ll run into. GGUF runs everywhere - Macs, mixed CPU/GPU setups, systems where the model doesn’t quite fit in VRAM. Universal format. EXL2 is NVIDIA GPU-only but faster when the model fits entirely in VRAM.
On Apple Silicon you want GGUF, via MLX or llama.cpp. Got enough NVIDIA VRAM? EXL2 gives you the best speed. And if you’re not sure, go GGUF - it works everywhere, and you can worry about speed once you know you like the model.
If you graduate to vLLM you’ll meet a third family - AWQ and the FP8 block formats - and that’s where quantisation support on consumer cards gets properly treacherous. I covered the trap that costs people an evening in VRAM traps in local LLM inference .
Hardware compared
| Hardware | VRAM | Price (GBP) | tok/s (14B) | tok/s (27B+) | Best For |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12GB | ~200 | ~42 (8B) / ~23 (14B) | - | Cheap experiments, 7-14B only |
| RX 9060 XT 16GB | 16GB | ~300 | ~25 est | - | Budget AMD, 14B at Q8 |
| RTX 3090 (used) | 24GB | 380-800 | 28-36 | ~28 | Best value for serious work |
| Mac Mini M4 Pro | 24GB unified | 1,399 | ~15-20 est | - | Silent macOS, 13B models |
| Framework Desktop | 128GB unified | ~1,970 | ~15-20 est | ~10-15 est | 70B mini PC, modular |
| Corsair AI Workstation 300 | 128GB unified | ~2,000 | ~15-20 est | ~10-15 est | 70B mini PC, Corsair build |
| GMKtec EVO-X2 | 96GB alloc | ~2,000-2,500 | ~15-20 est | ~10-15 est | 70B mini PC, first to market |
| RTX 4000 Ada | 20GB | 1,150 | ~20-25 est | - | Multi-GPU builds, low power |
| Dual 3090 (NVLink) | 48GB | 1,800-2,200 | 30+ | 15-20 | 70B models, prosumer |
| RTX 4090 | 24GB | 1,600-2,000 | ~40-50 est | ~35 est | Fastest 24GB option |
| RTX 5090 | 32GB | 2,900-3,500 | ~50+ est | ~40 est | Single-GPU 70B, no sharding |
| DGX Spark | 128GB unified | ~3,800 | ~15-20 est | ~10-15 est | 128GB NVIDIA ecosystem |
| Mac Studio M3 Ultra | 192GB unified | 5,999+ (falling) | ~10-15 est | ~8-12 est | 100B+ models, the discount play |
| Mac Studio M5 Ultra | 512GB unified | TBC (ships 22 Sept) | - | - | 200B+ MoE models, capacity king |
| Quad 3090 | 96GB | 3,000-3,500 | 100-180 | 26+ | Maximum VRAM on a budget |
Benchmarked figures from Digital Spaceport and Hardware Corner. Estimates marked ‘est’ from Gemini research and community reports. Mini PC tok/s varies significantly by model size and quantisation.
Corsair VENGEANCE a7500 AIR
- GPU RTX 5090 32GB
- RAM 192GB DDR5
- CPU Ryzen 9 9950X3D
- Storage 6TB NVMe
Corsair AI Workstation 300
- Memory Up to 128GB unified
- VRAM Up to 96GB allocatable
- Chip AMD Strix Halo
- Form factor 4.4L compact
What I’d buy
Had someone asked me this question two years ago I’d have said “whatever has the most VRAM under a grand.” My answer hasn’t really changed. Under £800, a used RTX 3090 is still the obvious play. Twenty-four gig of VRAM for under £800, NVLink ready for the inevitable second card, and enough capacity to run every model up to 34B at decent quantisation. Exactly where I started, and knowing what I know now, I’d make the same call.
If £800 is a stretch, my honest advice is to wait and save rather than buy below the line. An RTX 3060 will run 8B models respectably - the benchmark table up the page shows it - but you’ll outgrow it the first time a model you want won’t fit, and 24GB is where that stops happening for a good while. With used 3090 prices settling around £380-480, the wait is shorter than it used to be.
Between £800 and £3,000
Things have shifted in this bracket since I first wrote the article, and it's the mini PCs that did it. The big change: you no longer need a Mac or a four-card tower to get past 48GB. A whole class of 128GB unified-memory boxes now ships. If you want 70B from something you can hide on a shelf, the Framework Desktop (AMD Strix Halo, 128GB unified, ~£1,970) is where I'd look first - modular, repairable, and you can upgrade the board when the next chip lands rather than binning the machine - with the BOSGAME M5 as the pick if price beats brand for you. NVIDIA's DGX Spark (128GB, ~£3,200) is the same idea with a CUDA badge, if you want the Nvidia software stack out of the box. Already on macOS with smaller models? The Mac Mini M4 Pro still makes sense, though note Apple hiked UK prices hard in June 2026, so the high-RAM Macs are no longer the quiet bargain they were. And if you want to go down the same rabbit hole I did - dead-quiet density in a workstation - RTX 4000 Ada cards still do it, though I have since moved to a pair of modded 48GB 4090s on vLLM for the raw speed.
Above £3,000
The RTX 5090 remains the pick for raw speed if you can stomach the price (budget £3,250 minimum as of September 2026) - one card, one slot, 70B without any of the multi-GPU headaches. For NVIDIA’s take on unified memory, the DGX Spark at £3,800 gives you 128GB and the full CUDA stack. For maximum capacity on a budget, the quad 3090 build on AM4 gets you 96GB of VRAM at about £30 per gigabyte - ugly, loud, and you’d struggle to find anything with that much VRAM for less money. And for maximum capacity full stop, the Mac Studio M5 Ultra’s 512GB rewrites the ceiling when it ships on 22 September; if that’s tempting, keep an eye on M3 Ultra clearance prices too, because the discount play I predicted has finally arrived.
Buy the most VRAM you can afford. Pretty much everything else is secondary.
Related posts
Continue reading.
- Beginner's GuidesClaude Desktop System Requirements (2026): Minimum & Recommended Specs6 Jun 2026
- Beginner's GuidesClaude Code System Requirements (2026): Specs, Setup, Gotchas6 Jun 2026
- Beginner's GuidesA beginner's guide to Claude hooks31 May 2026
- ExplainerWhat is a Marketing Engineer (and do you need one)?31 Aug 2026
- AI WorkflowsDynamically updating infographics and content (so you don't have to)28 Aug 2026
- Case StudiesThe parts Shopify Plus B2B doesn't do (and how we built them)28 Aug 2026