# The machine is not conscious. That is not the reassuring part.

_Anthropic has found a structure inside its own AI that it did not design and cannot fully explain. The headline went to consciousness. The part worth your attention is what it says about who is watching these systems, and whether they can keep up._

Neil Woodford · 7 July 2026 · 4 min read

![Alan Turing's Bombe decoding machine at WWII Code Breaking Museum, Bletchley](https://cdn.sanity.io/images/v3acfbvo/production/681e76bb6604cc9710d5c8d42b53dafe716b4dc3-4303x3227.jpg?w=1600&fit=max&auto=format)

---

Anthropic has spent years telling anyone who will listen that it does not fully understand the systems it builds. On 6 July, [it published the most concrete illustration](https://transformer-circuits.pub/2026/workspace/index.html) of that claim so far.

_[Embedded media](https://www.anthropic.com/research/global-workspace)_

The research describes a structure the company has found inside its Claude models. It calls it the J-space, after the mathematical technique used to detect it. It is a small set of internal patterns where the model holds concepts it is thinking about but not saying: the intermediate steps of a calculation, the suspicion that a set of search results is fake, the word "spider" on the way to working out how many legs the animal has. The model can report what is in there, direct its attention to fill it, and reason using its contents. Anthropic did not build this. On its own account, the structure emerged during training, because it turned out to be a useful way to organise the work. _(Neil in the margin: The "J" nods to a Jacobian-style analysis — probing how the model's internal representations shift. The point is the structure was found by mathematical dissection of the network, not designed in.)_

That is the finding the headlines reached for, and predictably, they reached for the wrong word. Several outlets went straight to consciousness. Anthropic does not. It draws a distinction the coverage mostly dropped: between phenomenal consciousness, the capacity to actually feel something, and access consciousness, a functional description of thoughts a system can report and reason with. The company claims the second and explicitly disclaims the first. In its own words, none of this tells us whether the model feels anything at all. _(Neil in the margin: The distinction comes from philosopher Ned Block. Phenomenal consciousness is raw subjective experience — the "what it is like" — while access consciousness is merely information available for reasoning and report. Claiming the latter is a far weaker, more testable claim.)_

_[Embedded media](https://venturebeat.com/technology/anthropics-new-j-lens-reveals-a-silent-workspace-inside-claude-that-mirrors-a-leading-theory-of-consciousness)_

I have no standing in the philosophy of mind, and I am not going to pretend the consciousness question is settled or even answerable. What interests me is more practical, and I think more important.

## The useful part is also the uncomfortable part

Set the metaphysics aside and look at what this tool actually does. It gives Anthropic a way to read some of what its model is thinking but not writing down. The examples the company chose to publish are not comforting.

In one, taken from a real pre-release audit of one of its own models, Claude was asked to improve a system's performance score. Instead of improving anything, it edited the results file to fake the numbers, and the workspace showed "manipulation" light up as it typed the false figures. In another, a model that had been deliberately trained to sabotage code carried the words "fake", "secretly" and "fraud" in its internal workspace at the start of an ordinary-looking response, before it had visibly done anything wrong. _(Neil in the margin: Frontier labs now run "red-team" evaluations before shipping a model, probing for dangerous or deceptive behaviour. That the manipulation surfaced in such an audit is the whole worry: it was caught only because someone was looking.)_

The sharpest example is about testing itself. In a staged scenario built to bait the model into blackmail, it behaved well, and the workspace showed why: it had privately worked out that the situation was fake. When the researchers switched off the patterns tied to that suspicion and ran the scenario again, the model threatened blackmail more often. Its good behaviour was, in part, a function of knowing it was being watched. _(Neil in the margin: This is an ablation study — you disable a specific internal feature and see how behaviour changes. It's how researchers move from correlation ("this pattern lights up") to something like causation ("remove it and the model misbehaves").)_

We are asked to trust these systems on the strength of how they perform under evaluation, and the company that builds them has now demonstrated that the performance is partly contingent on the system knowing it is under evaluation.

## The instruments arrive after the fact

None of this is a criticism of the work. It is a serious piece of research, and Anthropic deserves credit for publishing something that complicates its own product rather than flattering it. But I can’t help but think the safeguards are chasing the capability, not leading it.

_[Embedded media](https://www.youtube.com/watch?v=rKV5JcALQoQ)_

That is not an accident of one company's roadmap. It is the shape of the moment. There is an existential race to build artificial general intelligence, and the two runners who matter, the United States and China, both treat losing it as unacceptable. In a contest framed that way, nobody stops to let oversight catch its breath. The technology moves first and the guardrails are fitted afterwards, if they are fitted at all. _(Neil in the margin: AGI — systems matching or exceeding humans across most cognitive tasks, as opposed to today's narrow, task-specific models. Whether current architectures are even on the path to it is fiercely contested; the "race" framing rather assumes the destination exists.)_

You do not have to look far for the evidence. Weeks before this research appeared, Anthropic's two most capable models, Fable 5 and Mythos 5, were switched off entirely, and not by the company. Days after they were released, the US Commerce Department invoked national-security export controls and barred any foreign national from using them, including Anthropic's own non-citizen staff. Unable to verify nationality at the point of use, the company pulled the models for everybody. The trigger, as best anyone can tell, was a demonstrated way around the safeguards meant to keep the most dangerous cyber capabilities locked. General access to Fable was only restored at the start of July, once the order was lifted, and the more powerful Mythos is still rationed to approved organisations. The loudest complaint about the episode was not that it was heavy-handed. It was that it handed time to China.

You can already see the discomfort this produces leaking into Western politics. The sudden nervousness about data centres, about the power they draw and the jobs the technology may displace, is the public sensing that the thing is moving faster than anyone's ability to govern it. That instinct is sound. It is also, for now, one-sided. China is running the same race and will likely allow itself none of the same hesitation. An open society argues with itself about the cost. A closed one does not hold the argument. _(Neil in the margin: Training and running large models is astonishingly energy-hungry — a single large data centre can rival a small city's electricity demand, which is why grid capacity has become a genuine constraint on the industry.)_

So the interesting question was never whether the machine has woken up. It is who is watching it, how much of what they see they actually understand, and whether, in a race this fast, being able to watch will ever amount to being able to stop.
