The thing I like about the Continuous Thought Machine is not that it is “brain-inspired.” That phrase has been stretched until it means almost nothing. What I like is much simpler: the CTM gives a neural network something like a private clock.
A transformer mostly thinks by writing. If it wants more computation, it usually has to emit more tokens, or pretend to reason in public through a Chain-of-Thought. The CTM instead has an internal temporal dimension: it can keep updating itself without necessarily producing more external symbols. This sounds small, but it changes the geometry of the problem. Compute is no longer glued to text length.
The central object is the synchrony matrix. Rather than treating all internal units as a static vector of activations, the CTM asks which units are moving together through time. The model’s state is not only “what is active?” but “what is co-varying?” That is a much richer kind of internal evidence.
If you have not seen the dynamics, the official interactive demo is worth playing with before reading further:
The annoying marriage of data and compute
Current language models have a strange habit: when they need to think longer, they often need to speak longer. This is useful engineering, but conceptually awkward.
- More reasoning often means more generated tokens.
- More generated tokens means more dependence on the distribution of written reasoning traces.
- Written reasoning traces are not the same thing as thinking.
This does not mean Chain-of-Thought is useless. It clearly works. But it also entangles two things that I would like to separate: the amount of external data consumed or produced and the amount of internal computation performed.
In humans, this distinction is obvious. You can stare at a chess position, look away, and continue working internally. You can understand a sentence, stop reading, and keep turning it over in your mind. Whether language is necessary for thought is a complicated question , but it is clearly not the whole story.
The CTM is interesting because it makes this separation natural. Its synchrony matrix is an internal computational object. It scales with the model’s latent dynamics, not directly with the length of the input sequence. So the question becomes: if the model has this private time, how should it use it?
Looking, then dwelling
When I started poking at the CTM, two behaviors bothered me.
First, the model often seemed to finish early. On CIFAR-10, much of the useful decision-making appeared to happen around the first handful of ticks. After that, additional time was not always additional thought.
Second, the model kept receiving fresh sensory input at every tick. This felt wrong. Not wrong in a moral sense, but wrong in the way a simulation of thinking feels wrong when the character never blinks.
A useful cognitive rhythm is not “look forever.” It is more like:
- sample the world,
- form a candidate internal picture,
- stop letting every new pixel bully the latent state,
- refine.
So I tried to make the CTM learn something closer to that rhythm.
A small gate for perception
The first modification was a Perceptual Gate. At each tick, the model produces a retention value . This value decides how much of the next state should come from the model’s current internal state versus the new observation :
When is close to 0, the model is still looking. When is close to 1, it is mostly dwelling internally.
I like this intervention because it is almost embarrassingly direct. There is no mystery variable called “reasoning.” There is just a knob that says how much the model trusts perception right now.
A nudge toward rhythm
The second modification was a small loss term. I wanted the model to discover a temporal shape: first explore, then consolidate.
I defined two moments:
- : the moment of high perceptual novelty.
- : the moment of large internal change.
Then I added:
This rewards low retention while the model is still looking, and high retention once it should be dwelling. It is not meant to force a rigid schedule. It is more like tapping the table in the tempo I want the model to hear.
What changed
I compared a baseline CTM against the combined version: Perceptual Gate plus the rhythm loss. The task was CIFAR-10, trained for 200k iterations.
Accuracy became less twitchy
The first thing I noticed was not just higher accuracy, but a different texture of learning. The modified model was smoother. The baseline had more of the usual nervous motion: up, down, up, down, as if each checkpoint was renegotiating the same fragile agreement.

But there was a cost. The modified model overfit heavily. This is not surprising: I gave it more machinery and a more opinionated training signal. It used both.

The decision moved later
The baseline CTM often commits early. It tends to become confident around tick 12. The modified model waits longer, with decisions spreading toward the middle ticks.
That delay is not automatically good. A model can also wait because it is confused. But paired with the accuracy curves, the shift looks more like verification than hesitation. The model reaches a plausible answer, then uses extra internal time to make it less brittle.

More time stopped being harmful
This was the most interesting result to me. In the baseline, later ticks can become actively worse. Letting the model “think” longer is not necessarily helpful; sometimes it just drifts.
The modified model does not show the same collapse. Its accuracy stays much flatter across time.

This matters because inference-time compute is only useful if extra computation remains coherent. A model that gets worse when allowed to continue is not thinking longer. It is dissolving.
Early exits improved
Because the modified model remains accurate across the trajectory, it also works better with early-exit rules. If we stop once confidence crosses a threshold, the gated model can reach that threshold quickly without paying as much in accuracy.

This is the nice version of inference-time scaling: not “always spend more compute,” but “spend compute until the internal state has settled.”
The gate learned the story
Finally, I looked at . If the whole setup was doing what I hoped, retention should start low and rise over time. That is what happened. The curve has the shape one would want: early openness to perception, then increasing internal closure.

Why I care about this
The point is not CIFAR-10. The point is that “thinking time” should be an object we can shape.
A lot of current scaling feels like increasing the size of the mouth: more tokens, longer traces, larger contexts. The CTM points toward a different axis. Give the model an internal temporal workspace, then ask how it should regulate contact with the world.
The Perceptual Gate is a small experiment in that direction. It is probably not the final mechanism. It may not even be the right mechanism. But it makes one thing visible: a model can learn when to stop looking and start dwelling.
That seems like a primitive worth taking seriously.
