For most of my life, I assumed everyone else was faking it.
When teachers said “picture an apple,” I thought it was just a metaphor, a shorthand for “think about an apple.” Only much later did I learn that most people really do get vivid inner images, complete with color and texture. There’s even a name for people like me who don’t: aphantasia. It’s usually described as having “no mind’s eye.” You understand the words, you can reason about shapes and scenes, but when you try to visualize, nothing shows up. It’s like typing DISPLAY IMAGE on a terminal with no monitor connected.
For decades, mental imagery research treated this inner cinema as essential. Strong spatial reasoning meant strong visual imagery. People like me were puzzles on the fringe, fascinating outliers who somehow managed without the supposedly fundamental tool.
Now large language models are kicking the door in. In September 2025, Morgan McCarty and Jorge Morales released a paper with a title that sounds written just for me: “Artificial Phantasia: Evidence for Propositional Reasoning-Based Mental Imagery in Large Language Models”. Their question is simple but provocative: Can a text-only model like GPT-5 solve classic “imagery” tasks as well as humans, even though it has no pictures in its head at all ?
The answer is yes. That has profound implications for how we think about aphantasia, AI capabilities, and what “thinking in images” really means.
The Finke Task Challenge
The experiment they used comes from Roger Finke’s 1989 research, a workhorse in mental imagery literature for decades. The task works like this: You start with a simple shape like a capital J. You mentally transform it through rotation, mirroring, or repositioning. You add another feature (a line, a dot). Then you identify what familiar object the result resembles.

Imagine a capital D, flip it 90 degrees to the left. The place it on top of capital J. Most people recognize an umbrella. Cognitive scientists historically used this as evidence that we store pictures in the head. The reasoning was: to perform multi-step transformations, you must manipulate quasi-visual representations. Language alone supposedly wasn’t enough.
That makes it a perfect stress-test for LLMs, because they are the opposite of the old theory. They only have language.
Text-Only Models Outperform Humans
McCarty and Morales ran a clever comparison study. They designed 60 multi-step transformation tasks, many brand-new to avoid training data contamination. They tested 100 human participants on all of them. Then they gave identical text-only instructions to models like o3 and GPT-5.
Human raters scored all responses to keep evaluation consistent. If the final answer matched the target object, it earned full credit. Near-misses received partial credit.

A system with no eyes, no visual cortex, and no pictures beat humans on a task designed to prove the importance of mental imagery. If you’ve grown up with aphantasia, that outcome feels oddly familiar. I’ve spent my whole life solving “visual” problems by leaning on language and logic. I don’t rotate a shape; I rotate a description of a shape. The models seem to be doing something similar.
When No Pictures Don’t Hurt
The authors didn’t stop at overall scores. They wanted to know how individual differences in imagery vividness affected performance. Participants completed the Vividness of Visual Imagery Questionnaire (VVIQ), a long-standing tool that asks you to rate how clearly you can picture different scenes.
If the traditional story held, people with rich, vivid imagery should dominate the task. People with aphantasia should struggle.
That’s not what happened. The sample showed a broad spread of VVIQ scores, from aphantasia to hyperphantasia. But the data revealed no positive correlation between imagery vividness and task performance. One participant with true aphantasia scored near the top of the entire group. Several hyper-vivid imagers performed below average.
This aligns with earlier aphantasia research. A 2022 study found that aphantasics show similar accuracy to controls on many imagery-linked tasks. The differences often appear in response time, not correctness. Lacking a mind’s eye doesn’t mean lacking the ability to complete the task. You might take longer or use different strategies, but the capacity remains.
The new twist from Artificial Phantasia is that large language models now sit on the same side of that argument. They’re extreme aphantasics with no visual representation at all, yet they solve the task and outperform typical humans. That forces a fundamental question: maybe the real engine behind these “imagery” tasks isn’t imagery at all.
Thinking With Sentences, Not Screenshots
In cognitive science, a long-running debate splits two camps. The pictorial view argues mental images are like inner pictures stored in a map-like format. The propositional view claims mental imagery is really structured descriptions: relationships, facts, and rules expressible in language.
The LLM results clearly support the propositional side. A text-only model can’t manipulate pixels inside its “head,” but it can maintain a structured state like: “Shape = capital J; orientation = vertical; curve at bottom.” “Rotate shape 90 degrees clockwise.” “Add horizontal line at top connecting both sides.” “Check which stored concept best matches this composite”.
That’s verbal scaffolding, not imagery. Many aphantasics describe doing exactly this, walking through language-like steps and labeling the result. The authors call this propositional reasoning-based mental imagery. The “imagery” lives in the relationships, not in any inner photograph.
For someone like me, that’s strangely liberating. It reframes aphantasia from “missing a basic cognitive tool” to “defaulting to a different representational format.” It isn’t better or worse, just another path up the same hill.
Reasoning Tokens as Cognitive Budget
One of the most important findings involves how the authors manipulated reasoning effort in GPT-5. They used a reasoning-style interface where you can specify “minimal,” “low,” “medium,” or “high” effort, basically controlling how many internal tokens the model burns before answering.
The pattern is remarkably clean. At minimal reasoning, GPT-5 falls below human performance. At low and medium, it climbs into human range. At high reasoning, it reaches the top scores, clearly beating the human average.
This matters enormously for AI governance and deployment. We often treat model quality as a single fixed property: “GPT-X is smart; model Y is dumb.” In reality, there’s a knob you can turn, the cognitive budget you allow the model to spend.
Starve it of tokens for quick responses, and you get shallow reasoning and more mistakes. Allow more internal steps, and it performs tasks that look suspiciously like advanced human cognition.
Moving Beyond Pictures
We still don’t know exactly how large language models represent intermediate states when solving these tasks. We don’t know whether human aphantasics and LLMs are “doing the same thing under the hood.” That’s a question for neuroscientists and computational modelers.
But we can say this much: Mental imagery tasks don’t require literal internal pictures. Aphantasia doesn’t equal cognitive deficit. It’s a different representational style. Large language models achieve human-like performance using propositional structures and sufficient reasoning tokens.
For educators, consultants, and AI Officers, the lesson is practical. Don’t obsess over whether AI “really sees” or whether people “really visualize.” Focus instead on what representations are available, how they can fail, and what tools we can layer on top (from prompts to policies to external visualizations) to support how different minds actually work.
Some of us never had a mind’s eye. Now we have something stranger: an artificial phantasia running on silicon, ready to think in our place or alongside us, one token at a time.
Further Reading
McCarty, M., & Morales, J. (2025). Artificial Phantasia: Evidence for Propositional Reasoning-Based Mental Imagery in Large Language Models. arXiv:2509.23108.
Pounder, Z. et al. (2022). Only minimal differences between individuals with congenital aphantasia and those with typical imagery. Cortex, 148, 180–192.
Marks, D. (1973). Visual imagery differences in the recall of pictures. British Journal of Psychology.





