Australian startup Springboards has released Flint, an LLM fine-tuned on Alibaba's Qwen 3 to give more varied answers to open-ended prompts than the mainstream chatbots. The pitch rests on a demo that is hard to unsee: ask ChatGPT, Claude, or Gemini for a random number between 1 and 10 and you will almost always get 7. Ask again and you get 3 or 4. Ask a third time and you get 8 or 9. Flint, in one live test, returned 3.7916.
The pattern is not a party trick. In November a team of researchers won the best-paper award at NeurIPS with a study titled "Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)." They prompted 25 different LLMs — including top US commercial models and open-source releases from China — 50 times each to write a metaphor about time. Most of the 1,250 responses were a version of "Time is a river" or "Time is a weaver." The convergence held not only within each model but across them.
Springboards cofounder and CEO Pip Bingemann has turned the finding into a sales demo. Asked to name a type of car, ChatGPT and Claude reach for a Toyota or a Honda; Flint returned a Ford F-150. Asked for a tagline for New Balance running shoes, Claude and ChatGPT both produced "Run your way." Flint gave "Built to last, run to win" — not a Cannes contender, but different. Ask ChatGPT to name a band and it will produce a list dominated by "glass," "neon," "velvet," and "static": in one test, 56 suggestions led by "Glass Harbor," with "Static Empire," "Neon Hearts," and "Velvet Echo" further down. Gemini's 15 suggestions included "Static Horizon."
Key facts
- 01Springboards built Flint on top of Alibaba's open-source Qwen 3 model, targeting variety in open-ended prompts rather than raw benchmark scores.
- 02A NeurIPS best-paper winner tested 25 LLMs, prompting each 50 times for a metaphor about time — most of the 1,250 responses were variants of "Time is a river."
- 03Ask ChatGPT, [Claude](/claude), or [Gemini](/gemini) for a random number between 1 and 10 and the answer is almost always 7; Flint returned 3.7916.
- 04ChatGPT produced 56 band-name suggestions led by "Glass Harbor"; Gemini gave 15 including "Static Horizon" — the same neon-glass-velvet-static cluster.
- 05Flint is aimed at advertisers and marketers via Springboards' brainstorming tool, which also routes prompts to ChatGPT and Claude.
The researchers behind the NeurIPS paper speculate the homogeneity comes from models being trained in similar ways on similar data for similar tasks. OpenAI told MIT Technology Review that training for reliable and coherent answers pushes models toward familiar, high-probability responses, and that pushing harder for novelty tends to make outputs weaker or less reliable. OpenAI also noted that the study looked at 2024-vintage models that have since been updated.
Springboards' Flint is built on Qwen 3 rather than a from-scratch foundation model. "We're a small team," says cofounder and CTO Kieran Browne. "Training a foundation model is not on the table for us. It's just too expensive." The company's product is a brainstorming tool for marketers and advertisers that lets users drag text from several models — including ChatGPT and Claude — into a workspace and recombine it. Flint slots in as the option users pick when they want the answer to veer off-piste.
The obvious first lever for variety is the temperature setting, and it was the first thing Springboards tried. It didn't work. Dialing OpenAI's temperature to its maximum produced responses that switched from English into code midway through a sentence. Springboards concluded that temperature is a blunt instrument — turning up randomness across every token wrecks coherence. Instead the team trained Flint to identify the specific points in an output where more variety is available (the destination in "Where should I go in Europe?") and inject the oddball only there.
Zoe Scaman, founder of business strategy shop Bodacious and chief strategy officer at 77X — the direct-to-fan platform set up by LA Lakers guard Luka Dončić — has been testing Flint against Claude, Gemini, and ChatGPT on a classic MBA prompt: reinvent a finance company for today's youth. The three mainstream models converged on financial-literacy edutainment. Flint suggested rebranding the concept of wealth accumulation itself. Scaman calls the prototype uneven — "it sometimes falls over when you start pushing it too far" — but says the premise is powerful.
Maximilian Weigl, cofounder and chief strategy officer at marketing firm Uncommon, uses Flint alongside ChatGPT, Claude, and Gemini. He also concedes the limits of the pitch: nine times out of 10, he says, the average is fine and mass-market familiar is what clients want. He also warns his team against copy-pasting from any model, Flint included.
“You can't really create something boundary-breaking with tools that pull you back to the average.”— Maximilian Weigl, Cofounder and chief strategy officer at Uncommon
The caveats matter. Flint sits atop a third-party open-source base, its variety trick is scoped to open-ended creative prompts, and the whole exercise assumes a user who actually wants a Buick answer instead of a Toyota one. For coding, research, and most enterprise workflows, the industry's convergence on high-probability answers is a feature, not a bug — reliability is what customers pay for, and the same training pressure that produces "Time is a river" also produces the SWE-bench gains that sell contracts.
Still, the homogeneity finding is a real product opening that the frontier labs have structural reasons to under-serve. Every incentive at OpenAI, Anthropic, and Google DeepMind points toward safer, more coherent, higher-probability outputs, because that is what enterprise buyers reward and what red-teamers demand. That leaves a gap on the creative end of the market — brainstorming, naming, positioning, campaign ideation — where the same conservatism reads as a flaw. Springboards is small and Flint is a prototype, but the wedge is genuine: if the next tier of AI differentiation is not raw capability but distribution of outputs, expect more startups to build on open weights and sell variety as the feature.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




