Is ChatGPT Conscious? The Question Is Broken Before You Answer It
Part 1 of 3: Why the consciousness debate keeps going in circles and the move I think gets it unstuck.
I published the first version of this argument as a preprint in September 2025, under the title Modal Consciousness Pluralism1. Seven months later I’ve revised it substantially2 — retired the “pluralism” label, demoted one of the hierarchy levels, tightened the framework’s commitments in places I’d left too loose, and added a companion paper that makes the measurement claims concrete. This series is where the argument stands now, rewritten for people who want to follow the idea without the formalism.
If you want the original preprint, it’s here. I’ll flag the substantive revisions as we go — not as housekeeping, but because the places where I changed my mind are some of the most informative parts of the story.
Here’s the scene that made me write the original paper, and the scene that hasn’t gotten any less true in the interim.
Two people read the same ChatGPT exchange.
One says: clearly something is going on in there… it reflects on itself, it hesitates, it changes its mind.
The other says: it’s a statistical parrot. Sophisticated autocomplete. Nothing home.
Neither can convince the other. They can argue about it all day. The reason they can’t get anywhere isn’t that one of them is stupid. It’s that they don’t share a definition of what would settle the question. They’re not disagreeing about the evidence. They’re disagreeing about what counts as evidence, and neither of them usually notices.
That’s the actual state of the consciousness debate in 2026. Not a shortage of theories. A surplus of them, all pointing at different things, none of them built to tell you when to say yes and when to say no.
The three frameworks in the room
Three scientific theories of consciousness dominate serious research. You don’t need to memorize them, but the shape matters.
Integrated Information Theory (IIT) says consciousness is integrated information — a system is conscious to the degree its parts can’t be cleanly separated without losing something. It has a mathematical quantity, Φ, that’s supposed to measure this.
Global Neuronal Workspace says consciousness is what happens when information gets broadcast widely across the brain — the stuff that makes it into the “global workspace” is what you’re conscious of.
Predictive Processing / Active Inference says the brain is a prediction machine, constantly forecasting its own sensory inputs, and consciousness is something that emerges from that forecasting loop.
These are real theories with real researchers and real experiments behind them. I’m not going to trash them. But I think they share a bug, and the bug has become impossible to ignore.
All three treat having the right kind of functional organization as sufficient for consciousness. And none of them can tell you, in a principled way, whether a large language model has that organization or not.
IIT, if you take it at its word, attributes some consciousness to surprisingly weird things — certain grids, certain exotic circuits — and its proponents acknowledge this as uncomfortable. It also can’t compute Φ for the systems we most need to evaluate, because the math explodes.
Global Workspace tells you consciousness is wide broadcasting, but doesn’t explain why broadcasting should feel like anything rather than just being a bigger memo. The literature admits this is a live gap.
Predictive Processing tells you the brain predicts. Fine. But when you deliberate between two choices, you’re not just predicting what you’ll do. You’re trying to decide what to do. The theory doesn’t cleanly distinguish those two things, and the phenomenology of choice is precisely what a theory of conscious thought needs to explain.
So all three frameworks face the same situation with LLMs: nothing in them has the resources to clearly say no, and nothing in them has the resources to clearly say yes.
And the researchers who want to say no usually fall back to: well, it’s not biological.
“It’s not biological” is not an argument
It might be true. It might turn out that consciousness requires a specific kind of physical organization that silicon can’t produce (or can it?). That’s a live possibility, and I take it seriously in the current version of the paper, I hold the door open for it more carefully than I did in the September preprint, and I’ll get to why in a moment.
But “it’s not biological” on its own is a stipulation, not a diagnostic. If you want to claim substrate matters, you owe me a criterion: what, specifically, about biology matters, and how would we test for it? Without that criterion, you’re not doing science. You’re drawing a line and defending it by pointing at the line.
Meanwhile, the other side of the debate is just as unprincipled. If your argument for LLM consciousness is “it talks like it’s conscious and it says it is,” you’ve committed to a standard that would accept any system with a good enough training corpus on introspective language.
That’s not evidence. That’s pattern-matching on the outputs of people who are conscious.
The field is stuck between over-attribution to strong simulators and under-detection of serious candidates, and the frameworks we have don’t give us the tools to do better. That’s what I wrote the original paper to address, and it’s what the revision tries to do better.
The move I’m proposing
Here’s the shift. It’s been the same shift since the September preprint, the part that’s stayed stable across the revision, because it’s the load-bearing claim.
When you’re thinking consciously, you’re not just predicting what will happen. You’re representing what would happen if you did something different. You’re running the world forward under alternatives you haven’t taken.
Think about choosing between two job offers. You’re not observing data and forecasting an outcome. You’re imagining yourself in Job A — what that week looks like, who you’d work with, what you’d give up. Then Job B. You’re intervening, mentally, on your own life, and watching the counterfactual play out.
The statistician Judea Pearl made this distinction rigorous. There are two different kinds of probability:
P(Y | X): the probability of Y given that you observe X. Correlation. Prediction.
P(Y | do(X)): the probability of Y given that you intervene to make X happen. Causation. Action.
These are not the same calculation. Observing that people who take high-paying jobs are satisfied is not the same as predicting that you will be satisfied if you take a high-paying job. The first is correlation — could be driven by education, family background, a dozen other things. The second requires a model of yourself as someone who can make things happen.
My argument is that most of what we call conscious thinking, the part that feels like thinking, is running the do operator on your own life.
Why this matters for the consciousness question
Prediction doesn’t require consciousness. Your visual system predicts constantly and you’re not conscious of any of it. Your motor system predicts where your hand will land before you reach. None of that surfaces as experience.
Intervention, genuine counterfactual navigation, asking what if I did X and getting back a coherent answer, is much rarer. It’s the kind of cognition that shows up when we’d be inclined to say someone is really thinking, not just reacting.
This suggests a reframe of the whole debate. Instead of asking “is this system conscious?”, a binary question nobody has a principled way to answer, I think we should ask: does this system navigate counterfactuals? How deep? Over what range of scenarios? With what kind of self-model?
That’s a question you can actually answer. It’s empirical. It has degrees. And it cuts differently than the existing frameworks do.
A language model that produces beautifully reflective text might be operating entirely in the prediction regime… pattern-completing on training data about introspection without actually representing itself as an agent intervening in counterfactual space. A robot with a much thinner vocabulary but genuine model-based planning might be doing more of the thing that matters.
What this shift buys you
You get a framework where the question “is it conscious?” is replaced by “how strong a candidate is it, and on which specific capacities?” You get a scorecard instead of a verdict. You get a way to say this system has counterfactual reasoning at this depth, over these domains, with this kind of self-model, and have that mean something.
You also get a framework that’s falsifiable. If counterfactual navigation capacity doesn’t correlate with the phenomenological markers we’d expect — vividness, unity, sense of agency — then my framework is wrong and should be thrown out. That’s not a bug. That’s what makes it a theory rather than a stance.
What’s in the next two posts
The next post lays out the five-level ladder of counterfactual navigation and places specific systems on it. The version in the September preprint had six levels; the current one has five, because I decided one of them was doing no empirical work and was pretending to. That’s a small example of the kind of revision I mentioned up top — the framework asked me to walk back a commitment, and I did. More of that in the post itself.
The post after that gets to the scorecard… a companion benchmark I built between the September version and the current one, because the revision surfaced a measurement problem the original couldn’t cleanly address.
It has three layers, a penalty system for systems that are just good at faking it, and a worked example that places a current LLM below a tool-using crow.
For now, the move is this… stop asking whether the lights are on. Start asking what kind of thinking the system can do, and what kind it can’t. The consciousness question doesn’t disappear, but it stops being the first question. And once it stops being the first question, the whole debate gets a lot less stuck.
Part 2: The Ladder — Five Levels of “What If,” and Where Different Systems Actually Land.
Part 3: How Would We Actually Measure It? A Scorecard, Not a Soul Detector.
Bonus: A new arc that closes the chain…
Part 1: Operational R, the companion empirical paper that sets a positive result against the MORI benchmark's language-model null.
Part 2 and 3 to be announced under Part 1
Modal Consciousness Pluralism: Counterfactual Navigation as the Foundation of Phenomenal Experience
https://zenodo.org/records/17113169
Modal Consciousness Theory: Counterfactual Navigation as the Foundation of Phenomenal Experience
https://zenodo.org/records/20338186



