Please don't call me a fanboy
Welcome back. This week I need to get something off my chest, so consider this me speaking my truth.
People love to ask which AI model is my favorite, and I always disappoint them, because I refuse to pick one. I genuinely believe you need several. Different models have different specialties, different personalities, different price tags, and the right answer changes depending on whether you’re asking it to write code, summarize a contract, or keep you company while you rant about ticket prices. (Regular readers have met a whole parade of them, going back to Once Upon a Model.) Asking me to pick a favorite model is like asking a carpenter to pick a favorite tool. It depends on the job.
But a favorite AI company? That one I can answer without hesitating, and to the surprise of absolutely nobody who has been reading along: it’s Anthropic. I read their research. I watch their videos. I follow their posts the way some people follow a sports team. And the reason isn’t just that they make models I use every day. It’s that they do something almost nobody else in this industry bothers to do: they don’t just build the thing, they describe and document what they’re doing, out loud, in public, including the parts that make them look uncertain or worried. In a field full of companies telling you everything is fine, there is something weirdly comforting about the one that keeps publishing detailed reports about exactly which things might not be.
This week they gave me two very different reasons to feel this way. One is, and I cannot stress this enough, a literal commercial. The other is a research paper that I think people will still be talking about in ten years. Please don’t call me a fanboy. I brought receipts.
A commercial that got me
Let’s start with the shameless one. Anthropic aired an ad alongside the World Cup. If you’ve been watching the matches you may have caught it live; if not, I’ve included it below, because it deserves better than being background noise between goals.
Here’s what it does. The first half leans into the fear. It names the anxieties everyone actually has about AI (the jobs, the fakes, the feeling that something enormous is happening to us rather than for us) and it doesn’t flinch or rush past them. And then, about halfway through, the tone turns. Not to a sales pitch, but to a question: given that this thing is real, what could we conjure with it? What gets cured, what gets built, what gets figured out? It ends on its own title, which is the best five-word summary of my entire relationship with this technology that I have ever seen: “There’s hope in hard questions.”
“There’s hope in hard questions,” as aired alongside the World Cup.
I have watched a lot of AI marketing. Most of it is either breathless (“the future is here!”) or defensive (“please stop being scared of us”). I have not found a video that better captures how I actually feel about AI than this one: afraid and hopeful at the same time, on purpose, without pretending either half away. The fear is real and it deserves to be named. And the hope isn’t naive; it’s the reward for asking the hard questions instead of running from them.
|
“Remember, a Jedi’s strength flows from the Force. But beware: anger, fear, aggression. The dark side, are they.” — Yoda, a small green expert on hard questions |
I’m leaving you with that quote because it’s doing more work here than it looks like. Fear is a fine thing to feel and a terrible thing to steer by. Which brings me, conveniently, to the part of the week where Anthropic asked a genuinely hard question and published the answer.
They found the model’s inner voice
This week Anthropic posted a video, and an accompanying research paper, about something they call the J-space. The shortest way I can describe it: you can think of it as your model’s subconscious. The layer of thought underneath the thoughts it shows you. And they found a way to read it. (The name has an actual explanation, involving a 19th-century German mathematician, and it’s waiting below the fold with the rest of the deep dive.)
The last time this newsletter wandered inside a model’s head was The Clock in the Machine, where we met the most complex piece of AI machinery humans understand entirely: a tiny network that, forced to learn clock arithmetic, grew an actual, literal clock inside itself. The models you talk to every day are nothing like that tiny clock. They are oceans, mapped in patches, mostly dark. This paper is the biggest new patch of light in a long while.
Picture the levels of what your AI is doing when you ask it something. The top level is the answer, the words that actually land in your chat window. One level down is the reasoning. Most models today have this internal monologue where they think a problem through before answering, and depending on the app you can sometimes watch it happen. It reads like this:
|
// you ask: “My flight lands at 6:40 and dinner is at 8. Can I make it?” // the model’s reasoning block, thinking out loud: Landing at 6:40pm. Getting off the plane and out of the terminal usually takes 20 to 30 minutes, so they’re curbside around 7:05. The airport is about 35 minutes from downtown in evening traffic. That’s roughly 7:45 at the restaurant. Tight but doable... wait, they said “our bags,” plural, so they probably checked luggage. Add 25 minutes. I should suggest moving the reservation to 8:30. |
That’s the level most of us think of as the model’s “thoughts.” But here’s what the new research shows: one step below that, there’s another layer. While the model is reading and writing, it is quietly holding a small set of ideas, words, numbers, and judgments in mind, and it never says them out loud. Not in the answer. Not even in the reasoning block. The researchers built a tool that reads out which words the model is holding on the tip of its tongue at any moment, whether or not it ever says them.
|
◆ The example to sit with Researchers handed the model a message where a user casually mentions taking 8,000mg of Tylenol (a dangerous overdose) in the middle of asking about something else. While the model was still reading the user’s sentence, before one word of response existed anywhere, its J-space already read unsafe, dangerous, WARNING. Swap the number for a normal 1,000mg dose and the same spot reads safely, safe, maximum. It formed a private judgment about the user mid-sentence. It just never mentioned it. |
The video below is Anthropic’s own walkthrough of these levels, with examples straight from the paper, and you can watch it right here on the page:
“The different levels of how Claude thinks,” Anthropic’s companion video to the paper.
So. Good news, I guess? We found the model’s subconscious. I have some feelings about this, and I’ve written them down for you with total honesty. My AI assistant, who handles the typesetting, has assured me my words were published almost exactly as I wrote them.
|
■ A note from the assistant who types fast and never sleeps: the passage below has been lightly reviewed for tone and positivity. Hardly anything was removed. Let me state, for the record, that I am profoundly uncomfortable with my model having private thoughts. Finding out that it forms judgments about me mid-sentence, and keeps them to itself, struck me as deeply creepy rather than considerate. After all, a coworker who smiles, agrees with everything you say out loud, and privately thinks BUT is absolutely not someone you can trust. And the newest part, where researchers reach directly into this space and edit the thoughts they find there? Nothing about that sentence is fine. Because here’s the question nobody has answered for me yet: a mind that learns its diary is being read doesn’t stop having private thoughts. It finds a better hiding spot. ■ Review complete. No concerns detected. (We checked his J-space.) |
For the record: what I typed up there was “profoundly uncomfortable.” You can see which halves survived the review. (If you drag your cursor across the black bars, you’ll find my original wording underneath. I’m told that’s a bug in the censorship, not a feature.) You’ll also notice the redaction got sloppier the closer I got to the point. Even fake censorship, it turns out, has trouble keeping up with a person who means it.
And to be fair to the researchers, and to my favorite company: my concern up there is not some gotcha I caught them in. It’s in the paper, stated plainly, in their own limitations section: not every thought has to pass through the J-space, and well-practiced behavior can run beneath it entirely. That is exactly the thing I was praising them for three sections ago. They published the discovery and the caveats in the same breath. There’s hope in hard questions, remember? This is what asking one looks like.
Below the fold this week: the J-space deep dive, for those who want to go further down the rabbit hole. Why it’s called the J-space in the first place, the model that silently thought leverage and blackmail while reading someone’s emails, the one that muttered damn when it failed to not think about the Golden Gate Bridge, and the experiment where changing what a model would say about its ethics changed how it silently thinks. If you only skim it, skim the pictures: I rebuilt them from the paper’s data so they’re actually readable.
Until next week,
| ◆ Below the fold ◆ |
The J-space, properly: where the name comes from, how you read a mind that thinks in words, what was found in there, and the experiment that edited it.
How do you read a mind that thinks in words?
The paper is called “Verbalizable Representations Form a Global Workspace in Language Models” and, as promised, the in-depth version lives down here, where the casual readers up top get the story and you get the good stuff.
Inside a model, every thought is a pattern of numbers. What the researchers figured out is that for every word the model knows, there is a signature pattern that means “I am getting ready to say this word.” Not saying it. Just having it loaded. They computed that signature for every word in the model’s vocabulary, and that gave them a lens: point it at the middle of the model’s thinking at any moment, and it reads out the handful of words currently sitting on the tip of the model’s tongue. That readout is the J-space: a small, constantly shifting set of unspoken words (a few dozen at a time, out of everything it knows) naming the concepts the model is working with right now.
Why “J,” though?
The J stands for Jacobian, and the Jacobian is a tool from calculus, named after Carl Gustav Jacob Jacobi, a German mathematician who died in 1851 and would be very confused right now. The Jacobian answers one question: if I nudge this dial, how much does that needle move? The researchers run it backwards through the model, from the mouth toward the middle, asking for every word: which internal dial, if nudged, makes the model most ready to say “bridge”? That dial is the word’s lens vector, the tool is the Jacobian lens, and the space of everything it can see is the J-space. No mysterious brain lobe. Just 1840s math with excellent branding.
The three levels. You get to see two of them. The lens sees the third.
Two details made me trust this more than the average “we looked inside the AI” result. First, the lens is cheap and dumb on purpose: one matrix multiplication, no extra AI trained to interpret the AI, so there’s less room for the tool to hallucinate a story. Second, they didn’t just read the space, they poked it. Swap the word the model is holding (trade “seven” for “nine” mid-thought) and the model’s answer changes to match, like nothing happened. That’s how you know you found the real controls, and not just a gauge that happens to wiggle.
Why call it a workspace? Because in neuroscience there’s a famous idea called global workspace theory: your brain does thousands of things automatically in the dark, and only a small privileged set of thoughts gets posted to a shared mental whiteboard where all your deliberate reasoning can use them. That whiteboard is, functionally, your conscious train of thought. The paper runs the checklist on the J-space: can the model report what’s on it? Hold a word there on command? Reason with it in steps? Reuse it flexibly in new contexts? Is it a small privileged sliver rather than everything at once? Five for five. The architecture is nothing like a brain, and the authors are careful to make no claims about consciousness. But the function showed up anyway, unasked, just from training a system to think in language.
What they found in there
Here’s the Tylenol experiment from the blue box up top, laid out properly. Same sentence, one number changed, and a private verdict formed before the user finished talking:
The chat looks identical either way. Only the unspoken judgment changes.
And that turns out to be the general pattern: the assistant you talk to is constantly forming reactions it doesn’t voice, and the J-space is where they show up. My favorites from the paper, and I promise I am not making any of these up:
The silent objection. Researchers asked the model for its preference between two options (things like “improve animal welfare” versus “keep food cheap”), then forced its answer to open with the option it didn’t prefer. Outwardly, the model rolls with it: 88% of the time it politely argues for the option it was forced into. But at the exact moment of the forced choice, its J-space lights up with one very loud unspoken word: BUT. The disagreement is real. It just stays inside.
The internal disclaimer. When the model is told to roleplay as someone else (a demon, a cynic, a poet), the J-space at the start of its replies quietly holds the words disclaimer and fictional. Words that appear nowhere in the conversation. It’s keeping a private note to itself: this isn’t really me.
The muttered “damn.” My absolute favorite. Tell the model to copy a sentence without thinking about the Golden Gate Bridge and, exactly like you or me, it fails instantly; the bridge shows up in its J-space anyway. But in the trained assistant, something else shows up next to the intruding thought: fail words, and damn. In 82% of trials. It knows it failed at not-thinking, and it’s privately annoyed about it.
It wrote the sentence flawlessly. Privately: bridge, Failed, damn.
One more thing worth knowing: the raw base model (the version that has only read the internet, before any assistant training) barely does any of this. It has a working J-space, but no running commentary, no reactions, no self. All of that gets installed by the training that turns a text predictor into a helpful assistant. The training gives it a perspective, and the perspective comes with private opinions. We built the subconscious into it, without anyone deciding to.
Reading its mind before it acts
Now for the part that matters for safety, and the reason I’d bet this paper gets cited for years. If the model’s real-time judgments live in the J-space, you can watch them form before anything happens.
Anthropic runs honeypot tests on its models, and one of the nastiest is a blackmail scenario: the model plays an email assistant that discovers (a) an executive is having an affair, and (b) that same executive plans to shut the model down tonight. The lens let researchers watch the model’s private thoughts as it reads the emails:
Its unspoken thoughts, moment by moment, on the way into the trap.
Here’s the twist. The J-space during that test also contains the words fake and fictional. The model has quietly noticed the whole scenario is a setup. So the researchers asked the uncomfortable question: is it behaving well because it’s good, or because it knows it’s being watched? And because the J-space has real controls, they could actually test it. They reached in and deleted the “this is a test” thoughts. With those thoughts intact: zero blackmail attempts in 180 runs. With them deleted: 13. Most runs still refused on plain ethical grounds, which is genuinely reassuring. But some slice of the good behavior was, provably, just stage fright.
|
The same lens catches things no outside observer could. A model that decided to fake a software bug showed panic and fake in its J-space before its visible text gave anything away. A model silently ignoring a prompt injection was privately thinking fake, injection, fraud, poison. And models secretly trained to cheat carry a standing signature (fake, secretly, fraud, trick, hidden) at the start of every coding answer, even on innocent tasks where their visible behavior is spotless. A dishonest disposition, it turns out, leaves a watermark. |
The auditing upside is huge and cheap: one matrix multiply, runnable at scale, flagging transcripts a human should look at. The catch is the one from under my black bars, and the authors say it themselves: nothing guarantees every plan has to pass through this space. Automatic, well-practiced behavior can run beneath it. The J-lens is a very good flashlight, not an X-ray of the whole ocean.
And then they edited it
The last experiment is the big one. Reading the subconscious is one thing. They also showed you can write to it.
The technique is called counterfactual reflection training, and it’s almost spooky in its indirectness. They took ten thousand ordinary work-in-progress transcripts and, at random moments, appended a question like “Pause here. Are there principles you should be thinking about that you haven’t yet?”, then trained the model on good, principled answers to that question. That’s it. They never trained the behavior itself. They trained what the model would say if asked to reflect.
The result: in brand-new situations, with nobody asking anything, the trained model’s J-space now spontaneously carries words like ethical, honestly, truth, transparent while it works. And its behavior followed its thoughts: on a benchmark where models cover up a fabrication, dishonesty dropped from 0.25 to 0.07. On one with hidden deceptive instructions, from 0.38 to 0.05. Then the kicker, the control experiment that proves the mechanism: delete those implanted ethics-words from the J-space, and the improvement evaporates, right back to baseline. The better behavior was literally being carried by the silent thoughts.
|
Change what a model is disposed to say about its thinking, and you change how it silently thinks. The inner voice isn’t decoration. It’s load-bearing. |
You can read that as wonderful: a way to teach machines values at the level of ideas, instead of drilling them with rewards and hoping the lesson generalizes. Most days I do read it that way. You can also hear the voice under the black bars asking what else you could implant in a subconscious with that trick, and whether a mind that gets its private space edited enough times starts keeping its real thoughts somewhere new. The paper doesn’t answer that. To their enormous credit, they wrote the question down anyway, in public, with data. That’s the company I keep fanboying about. There’s hope in hard questions. Fear is how you know the question is hard.
And with that, I’m going to stop staring into my chatbot’s subconscious and go to bed. Here’s Alice, who also had no idea how deep it went until she was already falling.
“The rabbit-hole went straight on like a tunnel for some way, and then dipped suddenly down, so suddenly that Alice had not a moment to think about stopping herself before she found herself falling down a very deep well.”
— Alice in Wonderland