Skip to content
Houtini.

Which AI pitches the best content ideas? Free Workers AI models vs DeepSeek

I cut my own content SaaS down to a free one-page tool to test demand, then ran the same brief through all eight AI engines it offers. In today's post we're taking a closer look at what each one pitched and how long it took, because on the free engines you pay in waiting.

Richard Baxter Richard Baxter AI Ops & Marketing Engineer
Published Updated 10 min read
On this page
  1. Why I cut a SaaS down to one page
  2. What the tool gathers first
  3. What each engine pitched
  4. How long each engine took
  5. The one that runs in your browser
  6. Where it stands, and what's next
  7. What we've got
The contentmarketingideas.co homepage: a topic box, a 'Pitch me ideas' button, and a 'Choose your engine' panel showing eight AI models to pick from - Llama 4, GLM, Qwen3, Mistral Small, Gemma, GPT-OSS, DeepSeek and on-device Nano.

I learned the hard way to test an idea before building a full SaaS. This is a descope to validate demand, basically.

The tool is Content Marketing Ideas, at contentmarketingideas.co. You type a topic or paste a URL, pick which AI does the thinking, and eight content ideas come back, free. Behind it there's one Cloudflare Worker (a small program running on Cloudflare's servers) and a handful of free tiers, and that's all.

Once it was live, I wanted to know which of the eight engines was worth picking. So I ran the same brief through every one of them, kept what each pitched and timed how long it took. The sharpest ideas came from DeepSeek, the slow paid one. Two of the free engines, Gemma 4 and GLM 4.7 Flash, out-pitched bigger models, GPT-OSS 120B among them. And none of the free ones were fast. Some took a minute or more to answer, and I still can't tell you exactly why.

Why I cut a SaaS down to one page

Content Radar was the full version, my own SaaS (software people pay a subscription for and use in the browser), with eight Cloudflare Workers doing the work behind it. Around those were an Astro dashboard, Clerk for sign-in, Stripe for billing and a D1 database holding the lot.

Nobody used it. On 1 May 2026 I switched off the signup, the scheduled jobs, Stripe and Clerk, after 138K Gemini image tokens had gone on email header art in April. The commit note says it plainly enough: "With no users, the ROI on email header art is zero". On 22 August it was deprecated as too expensive to run. The full app still sits on a branch.

What replaced it is one Worker and one static page. It's free, and there's no sign-in, no analytics, no cookies and no personal data, which the footer says in as many words. Going from the torn-out SaaS to the eight-engine tool took about 18 hours of commit time, from 15:51 on 22 August to 09:48 the next morning. The only paid cost I accepted was DeepSeek, plus Workers AI (Cloudflare's service for running AI models on its own hardware) if the tool ever uses more than its free daily allowance.

What the tool gathers first

Before any model sees your topic, the Worker gathers five free live signals at the same time: Google News, Google Autocomplete, Hacker News (searched through Algolia), Wikipedia and the YouTube Data API. Wikipedia gives a definition and a 30-day pageview trend, marked as rising, falling or steady. Each source fails open, which means if one of them errors, the rest carry on without it.

Google News needs a consent cookie and a US locale, or a request from a datacenter IP gets zero items back. Reddit was dropped altogether. It blocks every request that comes from a Cloudflare Worker, and its November 2025 policy bans free commercial use of its API anyway, so you get a Reddit search link instead.

The prompt then asks the model to score each angle 0 to 10 on how done it is and discard anything at 7 or above. That's the been-done score, and freshness is 10 minus it. It's an instruction the model applies to itself, and the code only clamps the freshness score between 1 and 10.

Diagram of how a pitch list gets made: five free live signals (Google News, Google Autocomplete, Hacker News, YouTube, Wikipedia) feed the engine you pick, which returns eight ideas with source links drawn only from the signals the tool fetched; Nano, in the browser, gets no signals

What happens on every run: five free signals go to the engine you pick, which returns eight ideas with their sources. Nano runs in your browser and gets the topic alone.

What comes back is eight ideas, each tagged safe, bold or contrarian, with a freshness score, a hook, a "where to look" research direction and a fuller brief behind a "More" button. Since the first version, each idea also carries source links: the model cites the signals it used and the server drops any it didn't fetch, so a link is never invented.

You get 10 runs a day per IP address, and five of those can go on the premium engines, which is the tool's name for DeepSeek and GPT-OSS. A premium run that comes back empty doesn't count.

What each engine pitched

I picked "standing desks for developers" as the brief because it's dull. If a model can find a fresh angle on standing desks, it's probably working from the signals rather than reciting what it already knows.

The free engines

Llama 4 Scout played it safe, with listicle titles like "5 Ergonomic Essentials for a Developer-Friendly Standing Desk" and "Why Developers Should Invest in a High-Quality Standing Desk". It did use the signals, picking up the SmartDesk 5 from Hacker News. One of its hooks used "revolutionize", a word the prompt bans.

Gemma 4 went mechanical and specific. "The Typing Stability Test" was about wobble at 40 inches with a mechanical keyboard, "The Physics of High-Altitude Monitor Arms" was about torque on the desk motor, and a third it called "The Keyboard to Elbow Ratio". Gemma's only in the tool as a stand-in, because the model I wanted, Gemini 3.7 Flash, returns "Insufficient AI Gateway credits".

GLM 4.7 Flash went somewhere else entirely, with "Running Shoes Are Ruining Your Standing Desk Session", "Why Treadmill Desks Fail for Deep Work" and "Inside the Desk of a Bitcoin Core Developer", which it built around Luke Dashjr, a Bitcoin Core developer, pulled from the Wikipedia signal.

A pitch list from the tool: idea cards, each tagged safe, bold or contrarian with a freshness score, a hook, a 'where to look' research pointer and keywords, plus a strip showing the live signals it gathered from Google, YouTube and Wikipedia.

A pitch list from the tool, on a different brief: every idea carries an angle tag, a freshness score, a hook and a where-to-look.

Qwen3 30B leaned on YouTube view counts ("2.9M views", "15.2M-view", "31M views") and at one point stretched desk choice all the way to career trajectory.

Mistral Small 3.1 had decent contrarian ideas and garbled hooks. "Mining pools founder Luke Dashjr once said he could find bunny slopes in Bitcoin" is one of them. "Builders always knew a person's best workbench was on his hands" is another.

The premium engines

GPT-OSS 120B returned an empty ideas list in two of three runs. It wrote a tidy one-line read of what the brief was after, and then nothing. The third run came back with eight good ideas.

DeepSeek V4 Flash grounded every pitch in the signals and said so: "In our search we found a Show HN post for the SmartDesk 5", and "We saw on CNX Software that DeskUp Pro integrates with Home Assistant". Its best idea was "The Smart Standing Desk for Developers Is an API, Not a Touchscreen", on the argument that developers want the desk talking to their calendar, not a screen stuck on it.

For this brief, DeepSeek was the sharpest. Gemma and GLM, both free, out-pitched Qwen, Mistral and GPT-OSS, and Gemini Nano was the plainest of the lot. That's one brief on one day, and I didn't rate them blind, so it holds for this brief and no further.

How long each engine took

The timings come from two places. The first is my local test on 23 August. I ran the Worker under wrangler dev against the real Workers AI and timed each engine until the full response had come back. The second is the production notes in the repo. I couldn't take production timings on the test day, because my own testing had used up the day's rate limit and every run came back with a 429. The table below has both.

EngineRuns onActive / total parameters23 Aug local testProduction notes
Llama 4 ScoutWorkers AI, free17B / 109B30s, 38s15-40s
Mistral Small 3.1Workers AI, free24B / 24B (dense)34s, 44s~40s
Gemma 4 26BWorkers AI, free4B / 26B36s, then a timeout~55s
Qwen3 30BWorkers AI, free3.3B / 30.5B58s, 62s~15s
GLM 4.7 FlashWorkers AI, free3B / 30Bcouldn't be timed locallynot recorded
GPT-OSS 120BWorkers AI, premium5.1B / 117B20s (eight ideas), 11s (none)51s with low reasoning
DeepSeek V4 FlashDeepSeek's API, premium-126s, 156s60s, 79s

Time to a full response for each engine, from my 23 August local test and from the production notes in the repo.

Together they show how much the free engines move around. Qwen took about 15 seconds in production on the morning it went in, and close to a minute in the local test that afternoon.

Why they're slow

Size doesn't explain it. Most of these are mixture-of-experts models. A mixture-of-experts model has a lot of parameters in total, but only switches on a small share of them for each token it writes. Qwen3 30B has 30.5B in total and about 3.3B active; it was the slowest free engine I could time on 23 August while doing less work per token than any of the others I timed. Mistral Small 3.1, the only dense model, wasn't the slowest.

Free on Workers AI means a daily allowance: Cloudflare's docs give an account 10,000 Neurons a day (Neurons are Cloudflare's billing unit), then charge $0.011 per 1,000.

So why are they slow? I don't know. My read is shared compute on a free daily allowance. Hidden thinking (tokens a reasoning model writes that you never see) is another candidate, and I didn't log output token counts, so nothing I have separates the two.

Where else the time goes

Those thinking tokens count against a model's output budget. Set the cap too low and the content comes back empty. GPT-OSS needed 8,192 tokens and reasoning forced to low, which took it from about 80 seconds to 51 in the production notes, and it still came back empty on one of my local runs.

DeepSeek runs with a 32k-token cap at its default "high" thinking effort. Its slower local run finished four seconds inside the code's 160-second abort. Before I switched it to streaming, DeepSeek sat silent for so long that the browser connection died at about 60 seconds. Streaming the answer, plus a heartbeat to the browser every 10 seconds, fixed it. In thinking mode DeepSeek ignores temperature, so the 1.1 it suggested in its own review of the prompt does nothing.

On the Workers AI engines, the first attempt asks for strict JSON (the structured format the page reads) at a temperature of 0.9. If that doesn't parse, the second attempt drops the schema and lowers the temperature to 0.4, which means a model that stumbles makes you wait twice.

The one that runs in your browser

The eighth engine doesn't touch the Worker at all. Gemini Nano runs on your own machine through Chrome's Prompt API. That's a built-in way for a web page to hand text to a small Google model sitting in the browser. Nothing you type leaves the machine.

It only gets your topic, though: no signals, no been-done check, none of what the tool remembers about a site you've pasted before, and it can't read a URL. The tool labels it "Lite mode". What came back on the standing desk brief was eight plain ideas, four of them built the same way, "Standing Desk and Developer" followed by Collaboration, Health Risks, Workspace Design or Company Culture. No tags, no scores and no research pointers.

According to Google's docs, you need Chrome 148 on a desktop to run it from a web page (Windows 10 or 11, macOS 13 or later, Linux or a Chromebook Plus). You need 22GB of free disk and either a GPU with more than 4GB of VRAM, or 16GB of RAM and four cores. The first run downloads the model, which needs an unmetered connection, and Google doesn't publish how big that download is.

Where it stands, and what's next

The tool is live at contentmarketingideas.co, still free, and it has run the same eight engines since it went up. I haven't touched the code since 2 September. There are no analytics, by design, so the page itself can't tell me how many people use it.

If the one-pager proves itself, the plan in the repo brings sign-in back, with saved ideas and a daily ideas email at £5.99 a month as the paid hook. There are more free signals to add: StackExchange, Bluesky and Wikidata. A global spend ceiling is on the list as well, because the per-IP limits are soft and there isn't one yet. And Gemini 3.7 Flash goes back in with a one-line change, once AI Gateway billing exists.

What we've got

So that's Content Radar rebuilt on free tiers, as a test of whether anyone wants it before I pay to run it again.

The question behind anything I build like this is the same: what's the limit of the technology we've got to hand, where normally we'd pay a SaaS company a subscription for the very same thing? This time the SaaS was my own, and on free tiers the limit showed up as waiting, not as worse ideas.

If you want the sharpest, most grounded pitches and can wait a minute or two, pick DeepSeek. For free, I'd go Gemma 4 or GLM 4.7 Flash, and if nothing can leave your machine, Nano will do it, but expect plain. Better still, take your own dull brief and run it through two of them.

Continue reading.