LLMs Fail Stroop Test as List Length Increases

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Suketu Patel led the study that applied the Stroop task to top large language models, testing their attention and executive control.
- GPT-4o scored 91% accuracy on five‑word lists, 57% on ten‑word lists, and fell to 15% on forty‑word lists, showing a sharp performance drop as list length increased.
- Claude 3.5 Sonnet maintained stable accuracy up to twenty‑word lists but dropped to 24% on forty‑word lists.
- GPT-5 showed similar declines as Claude Opus 4.1 and Gemini 2.5, failing to sustain high accuracy on longer or mixed lists.
- Gemini 2.5 and other models dropped to near‑zero accuracy on mismatched color‑word items, indicating they defaulted to reading the word rather than naming the ink color.
Why it matters: Developers of AI systems and users seeking reliable, distraction‑free performance now face a clear limitation: current large language models cannot consistently maintain task focus, especially in longer, conflicting sequences, reducing their utility for attention‑critical applications. This weakness may hinder adoption in fields like real‑time decision‑making, complex data analysis, and multi‑step instruction following, where sustained attention is essential.



