Skip to content
Houtini.
Contact
AI Tools ·23 August 2026

Eight AI models, one content brief: what each pitched (and why the free ones crawl)

Discuss and expand Ask ChatGPT Email LinkedIn

I built a free tool that pitches eight content ideas, and it lets you choose which AI does the thinking. So I ran the same brief through all eight models to see how differently they'd think. The ideas surprised me. The speed surprised me more.

The contentmarketingideas.co homepage: a topic box, a 'Pitch me ideas' button, and a 'Choose your engine' panel showing eight AI models to pick from - Llama 4, GLM, Qwen3, Mistral Small, Gemma, GPT-OSS, DeepSeek and on-device Nano.

I built a little tool called Content Marketing Ideas. You type in a topic, or paste your site's address, and it pitches you eight things you could write next, each one with a hook, an angle, a freshness score, and a note on where to go and research it. It's free. And the bit I'm most pleased with (probably more than I should be) is that you get to pick which AI does the thinking.

There are eight models sitting behind it, from a tiny fast one that runs on Cloudflare's network to DeepSeek's full reasoning model to a little thing that runs entirely on your own laptop. So I did the obvious thing and ran the same brief through every one of them, to see how differently they'd each pitch it. The brief was "standing desks for developers", dull on purpose, because a boring topic is exactly where you find out whether a model can think of an angle or is just padding. Here's what came back. And here's the thing nobody tells you when they sell you a free AI model: the fast ones aren't fast at all.

What you're choosing between

Before it asks a single model for anything, the tool goes and reads the room. It pulls what's being published right now from Google News, what people are typing into Google's autocomplete, what's being argued about on Hacker News, the top YouTube videos by view count, and a bit of Wikipedia context. Then it runs a "been-done" check: every angle gets scored 0 to 10 on how tired it is, and anything scoring seven or higher gets binned before you ever see it. That's the raw material. The model's job is to read all of that and pitch you eight things worth writing. (It's the same job I lean on a stack of MCP servers for content marketing to do inside Claude: work out what's worth writing before you write it. This just does it in one page, for free.)

So the model you pick is the editor, and they've all got different instincts. The free ones run on Cloudflare's Workers AI: Llama 4, GLM, Qwen3, Mistral Small, Gemma. Then there are two premium ones that think harder, GPT-OSS and DeepSeek. And one outlier, Nano, that skips the cloud entirely and runs in your browser (more on that one later, because it plays by completely different rules).

Same brief, same signals, same instructions. The only thing that changes is the mind reading it.

Same brief, eight different minds

A pitch list from the tool: eight idea cards, each tagged safe, bold or contrarian with a freshness score, a hook, a 'where to look' research pointer and keywords, plus a strip showing the live signals it gathered from Google, YouTube and Wikipedia.

What a pitch list looks like: angle tag, freshness score, hook and a where-to-look on every idea. (Different brief in the shot, same format.)

Llama 4 is the default, and it's the one you'd call safe. It gave me "5 Ergonomic Essentials for a Developer-Friendly Standing Desk" and "Why Developers Should Invest in a High-Quality Standing Desk". Competent, and it did use the live signals (it spotted the SmartDesk 5 that was doing the rounds on Hacker News), but these are the titles a content mill would hand you. One of its hooks reached for the word "revolutionize", which the instructions explicitly ban, and the smaller model just sailed straight past that rule. Fine for a quick spray of ideas. Not a lot of edge.

Then Gemma surprised me, and I don't think it's the model anyone would bet on. It went mechanical and specific: "The Typing Stability Test" (how much a desk wobbles at 40 inches when you're hammering a mechanical keyboard), "The Physics of High-Altitude Monitor Arms" (the torque a heavy dual-monitor setup puts on the desk motor), "The Keyboard to Elbow Ratio". Those are angles an actual desk-obsessive would pitch. GLM was in the same league, going for "Running Shoes Are Ruining Your Standing Desk Session" and, my favourite bit of resourcefulness, "Inside the Desk of a Bitcoin Core Developer", where it grabbed Luke Dashjr straight out of the Wikipedia snippet the tool had fed it and built a whole angle around him.

Qwen tried hard and slightly overdid it. It's the model that read the YouTube signals and couldn't stop citing them: "the 2.9M-view video", "the 15.2M-view desk build", "31M views". Thorough, but it leans on the numbers like a student padding a word count, and one or two of its angles are a stretch (it tried to connect your choice of desk to your career trajectory, which, no).

Mistral is where it got interesting in a different way. The ideas underneath were fine, some good contrarian ones about not over-complicating your setup, but the hooks came out garbled. One of them opened with "Mining pools founder Luke Dashjr once said he could find bunny slopes in Bitcoin", which is not a thing anyone said, because it isn't a sentence. Another: "Builders always knew a person's best workbench was on his hands." That's a mid-size model losing the thread mid-sentence, and it's worth seeing, because it's the exact failure the marketing never shows you.

GPT-OSS is the big one, a 120-billion-parameter model built to reason hard, and it's the flakiest of the bunch. Run it a few times and about half of them it reads the brief, writes a tidy summary of what my site's about, and then hands back an empty list. Nothing. The rest of the time it's perfectly good. That empty list isn't random, and the reason for it is the same reason the next section exists, so hold that thought.

DeepSeek is the one you wait for, and it earns the wait. It's the only premium model I let reason at full stretch, and it read the brief like an editor who'd done the homework, grounding every pitch in the actual signals the tool had gathered and telling you so as it went: "In our search we found a Show HN post for the SmartDesk 5", "We saw on CNX Software that DeskUp Pro integrates with Home Assistant." The best of them was "The Smart Standing Desk for Developers Is an API, Not a Touchscreen", which has a real argument underneath it: developers don't want a screen bolted to the desk, they want the desk to talk to their calendar. It also took the thick end of two and a half minutes to write that. Hold that thought as well.

So, for this brief: the sharpest ideas did come from DeepSeek, the slow expensive one, which is roughly what you'd hope. The surprise was underneath it. Gemma and GLM, both free, both models nobody would pick out of a line-up, out-pitched Qwen and Mistral and the bigger, pricier GPT-OSS, and they weren't miles off DeepSeek either. That won't hold for every topic. But it held for this one, and it's the whole reason the tool lets you switch.

Why the free, fast ones crawl

Here's the part that caught me out, and I built the thing.

None of these are fast. I timed them running locally against the real models, reading each response from the first byte to the last. Here's how long eight short ideas took:

ModelRuns onReasoningTime to eight ideas
Llama 4 Scout (17B)Workers AI, freenone30-38s
Mistral Small (24B)Workers AI, freenone33-44s
Gemma 4 (26B)Workers AI, freenone~36s
Qwen3 (30B)Workers AI, freenone58-62s
GPT-OSS (120B)Workers AI, premiumlow (forced)~20s*
DeepSeek V4DeepSeek's API, premiumfull126-156s

*GPT-OSS hands back nothing about half the time, and the quick runs are the ones where it gave up early. (GLM isn't in the table because my local test couldn't hold a clean connection to it long enough to time; on the live site it sits in with the other free ones.)

A 17-billion-parameter model taking half a minute to write eight short ideas is slow, and you can watch it get slower as the models get bigger: Qwen3 at 30B took a full minute. On my own rig, two RTX 4090s in the cupboard, a model Llama's size answers in under a second. So what's the difference? Compute. Cloudflare's Workers AI keeps these models free by running each request on a very small slice of GPU, and that trade is the whole game: you pay nothing, and you wait. Cheap and slow are the same coin. (If you want the other end of that trade, running a model on hardware you own so it answers instantly, that's a rabbit hole I've been down: best GPUs for running local LLMs .)

The two premium models are a different story, and reasoning is the reason. Both of them think before they answer, generating a pile of hidden "thinking" tokens you never see, and those count against the budget. That's why GPT-OSS comes back empty half the time: it spends its allowance thinking and has nothing left to write the ideas with. I've hobbled its reasoning to "low" on purpose, in two places, which is why the times it does work it's quick, about 20 seconds. DeepSeek I've left to reason fully, because it's the sharpest of the lot and worth waiting for, and it shows in the clock: two to two and a half minutes, every single time. It's the long pole, and I've made my peace with it.

One more thing that adds to it: if a model's first answer comes back malformed, the tool quietly asks it again with stricter settings. That's a safety net for the smaller models that occasionally return broken JSON, and it works, but a second attempt means a second wait. So a flaky model isn't just unreliable, it's twice as slow when it stumbles.

None of this is a complaint. Free models running on shared compute are a good deal, and I use them constantly. It just isn't "fast and free". It's free, and you'll wait a bit. Pick one.

The one that runs on your laptop

Nano is the odd one out, and the reason it exists is privacy. It runs entirely in your browser, on your own machine, so your topic never leaves your laptop. No API call, no server, nothing logged. If that matters to you, it's the only one here that clears the bar.

The trade is real, though, and you feel it in the output. Nano gets none of the live signals, none of the news, no been-done check, because all of that needs the network it's deliberately not using. So it pitches from what it already knows, and it shows: eight ideas, plainer titles, and four of them the same shape ("Standing Desk and Developer Collaboration", "...and Developer Health Risks", "...and Developer Workspace Design", "...and Developer Company Culture"). No angle tags, no freshness scores, no research pointers. It's also slower to start the first time, because it has to warm the model up in your browser before it can begin.

So it's blander, and that's not the model being weak, it's the model working blind. Private and offline, or sharp and signal-fed. You don't get both, and Nano is honest about which one it is.

So which one should you pick?

It depends what you're after, which is exactly why I didn't just hard-wire one in.

If you want a quick pile of ideas and don't mind a 30-second wait, the free ones are good, and Gemma and GLM are the two I'd reach for first after this test. If you want the sharpest, most defensible angles and you'll wait a bit longer for them, DeepSeek is the editor of the group. If you'd rather nothing left your machine, Nano, with your eyes open about what it can't see. And if you just want to feel the difference for yourself, run a boring brief through two of them back to back. That's when it clicks.

It's free, there's no login, and you can go and do exactly what I did right now: pick a model and pitch yourself eight ideas . Same brief, eight minds. Go and disagree with them.

By email

Get new posts by email.

Drop your email below and we will send you the next article when it lands. No spam, unsubscribe anytime.

More like this

Continue reading.

How to Do a Technical SEO Audit with Claude
AI Tools

How to Do a Technical SEO Audit with Claude

A free, step-by-step technical SEO audit with Claude: your Search Console history and a first-party crawl merged in one local database, ranked by recoverable clicks - the Screaming Frog alternative you run by conversation.

The VRAM traps: why a 16GB model wouldn't load on a 48GB card
Local AI

The VRAM traps: why a 16GB model wouldn't load on a 48GB card

A 16GB model would not load on my 48GB card, and the reason was a shortfall of one kilobyte in a memory no spec sheet mentions. These are the fit traps a VRAM figure will never warn you about.

The Dual-4090 96GB vLLM Benchmark & Runbook
AI Tools

The Dual-4090 96GB vLLM Benchmark & Runbook

Can a £6,200 modified Ada rig match enterprise MoE throughput? The living measurement record for a dual RTX 4090 48GB vLLM rig - every number measured here.

The broken rules of local LLM inference
Local AI

The broken rules of local LLM inference

I used to lock the clocks on my mining GPUs. The same instinct just helped kill five rules of local LLM inference on a £6,200, 96GB rig.

How to Stop MCP Servers Eating Your PC: One Docker Gateway for Every Claude Client
How-to Guides

How to Stop MCP Servers Eating Your PC: One Docker Gateway for Every Claude Client

My MCP list grew until orphaned node.exe were quietly eating a 128GB workstation by mid-afternoon. Here's how I put every server behind one Docker gateway - node on bare metal gone, secrets in one gitignored file, and a single URL every Claude client points at. One evening's work you'll feel every day after.

How to Use the Gemini API (and Why I Run It Next to Claude)
How-to Guides

How to Use the Gemini API (and Why I Run It Next to Claude)

Get a Gemini API key, make your first call in curl and Python, dodge the thinking-token trap that returns an empty answer, and see why running Gemini next to Claude is the real unlock. Written from production - and the bills.