Springboards' Flint Targets LLM Groupthink

SkimNews Take
Mainstream LLMs' identical answers likely reflect shared training data and alignment incentives that reward safe, consensus-friendly responses—so genuine output diversity is a problem only an outsider with no incumbent user base to protect would prioritize solving.
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Springboards built Flint, an LLM designed to counter the 'groupthink' tendency of mainstream models — when asked for a random number 1–10, Claude, ChatGPT, and Gemini almost always answer 7, while Flint returned 3.7916 in a fresh session
- A November research paper titled 'Artificial Hivemind' found 25 different LLMs converged on similar answers to open-ended prompts (1,250 metaphor responses about time mostly said 'Time is a river' or 'Time is a weaver') and won best paper at NeurIPS
- Flint runs on Qwen 3, an open-source model from Chinese tech giant Alibaba, because training a foundation model from scratch 'is not on the table' for the small Springboards team
- Instead of cranking up 'temperature' (which makes models incoherent), Flint selectively boosts randomness at specific points in its output — e.g., just before naming a European travel destination, not across the whole response
- Zoe Scaman (Bodacious, 77X) tested Flint against Claude, Gemini, and ChatGPT on a classic MBA case and said Flint suggested rebranding wealth accumulation rather than the standard 'teach financial literacy in a fun and funky way' answers
- OpenAI responded that training for reliable answers pushes models toward high-probability responses, and that the 'Artificial Hivemind' paper studied 2024 models that have since been updated
Why it matters: Every ChatGPT and Claude user is essentially getting the same answers, making the 'personal' chatbot conversation feel more like an echo chamber — a problem for creative professionals whose work risks becoming generic and undifferentiated. Springboards' workaround runs on Alibaba's Qwen 3, showing how third parties can layer novelty on top of open-source Chinese foundation models rather than training their own.

