Decoupling Compute and Data: Some Examples from the CTM

Decoupling Compute and Data: Some Examples from the CTM

January 24, 2026v1.0

The thing I like about the Continuous Thought Machine is not that it is “brain-inspired.” That phrase has been stretched until it means almost nothing. What I like is much simpler: the CTM gives a neural network something like a private clock.

A transformer mostly thinks by writing. If it wants more computation, it usually has to emit more tokens, or pretend to reason in public through a Chain-of-Thought. The CTM instead has an internal temporal dimension: it can keep updating itself without necessarily producing more external symbols. This sounds small, but it changes the geometry of the problem. Compute is no longer glued to text length.

The central object is the synchrony matrix. Rather than treating all internal units as a static vector of activations, the CTM asks which units are moving together through time. The model’s state is not only “what is active?” but “what is co-varying?” That is a much richer kind of internal evidence.

If you have not seen the dynamics, the official interactive demo is worth playing with before reading further:

The annoying marriage of data and compute

Current language models have a strange habit: when they need to think longer, they often need to speak longer. This is useful engineering, but conceptually awkward.

  • More reasoning often means more generated tokens.
  • More generated tokens means more dependence on the distribution of written reasoning traces.
  • Written reasoning traces are not the same thing as thinking.

This does not mean Chain-of-Thought is useless. It clearly works. But it also entangles two things that I would like to separate: the amount of external data consumed or produced and the amount of internal computation performed.

In humans, this distinction is obvious. You can stare at a chess position, look away, and continue working internally. You can understand a sentence, stop reading, and keep turning it over in your mind. Whether language is necessary for thought is a complicated question , but it is clearly not the whole story.

The CTM is interesting because it makes this separation natural. Its synchrony matrix S∈RD×DS \in \mathbb{R}^{D \times D} is an internal computational object. It scales with the model’s latent dynamics, not directly with the length of the input sequence. So the question becomes: if the model has this private time, how should it use it?

Looking, then dwelling

When I started poking at the CTM, two behaviors bothered me.

First, the model often seemed to finish early. On CIFAR-10, much of the useful decision-making appeared to happen around the first handful of ticks. After that, additional time was not always additional thought.

Second, the model kept receiving fresh sensory input at every tick. This felt wrong. Not wrong in a moral sense, but wrong in the way a simulation of thinking feels wrong when the character never blinks.

A useful cognitive rhythm is not “look forever.” It is more like:

  1. sample the world,
  2. form a candidate internal picture,
  3. stop letting every new pixel bully the latent state,
  4. refine.

So I tried to make the CTM learn something closer to that rhythm.

A small gate for perception

The first modification was a Perceptual Gate. At each tick, the model produces a retention value rt∈[0,1]r_t \in [0,1]. This value decides how much of the next state should come from the model’s current internal state ztz_t versus the new observation oto_t:

xt+1=rtzt+(1−rt)otx_{t+1} = r_t z_t + (1 - r_t) o_t

When rtr_t is close to 0, the model is still looking. When rtr_t is close to 1, it is mostly dwelling internally.

I like this intervention because it is almost embarrassingly direct. There is no mystery variable called “reasoning.” There is just a knob that says how much the model trusts perception right now.

A nudge toward rhythm

The second modification was a small loss term. I wanted the model to discover a temporal shape: first explore, then consolidate.

I defined two moments:

  • tlookt_{look}: the moment of high perceptual novelty.
  • tdwellt_{dwell}: the moment of large internal change.

Then I added:

λgate⋅[rlook2+(1−rdwell)2]\lambda_{gate} \cdot [ r_{look}^2 + (1 - r_{dwell})^2 ]

This rewards low retention while the model is still looking, and high retention once it should be dwelling. It is not meant to force a rigid schedule. It is more like tapping the table in the tempo I want the model to hear.

What changed

I compared a baseline CTM against the combined version: Perceptual Gate plus the rhythm loss. The task was CIFAR-10, trained for 200k iterations.

Accuracy became less twitchy

The first thing I noticed was not just higher accuracy, but a different texture of learning. The modified model was smoother. The baseline had more of the usual nervous motion: up, down, up, down, as if each checkpoint was renegotiating the same fragile agreement.

Test Accuracy vs Baseline
Figure 1. Test accuracy. The combined model is less jittery and reaches better peaks.

But there was a cost. The modified model overfit heavily. This is not surprising: I gave it more machinery and a more opinionated training signal. It used both.

Train Accuracy vs Baseline
Figure 2. Training accuracy. The same intervention that stabilizes test behavior also makes memorization easier.

The decision moved later

The baseline CTM often commits early. It tends to become confident around tick 12. The modified model waits longer, with decisions spreading toward the middle ticks.

That delay is not automatically good. A model can also wait because it is confused. But paired with the accuracy curves, the shift looks more like verification than hesitation. The model reaches a plausible answer, then uses extra internal time to make it less brittle.

Tick Distribution
Figure 3. Certainty over ticks. The baseline is impatient; the modified model distributes commitment across later internal time.

More time stopped being harmful

This was the most interesting result to me. In the baseline, later ticks can become actively worse. Letting the model “think” longer is not necessarily helpful; sometimes it just drifts.

The modified model does not show the same collapse. Its accuracy stays much flatter across time.

Per Tick Accuracy
Figure 4. Per-tick accuracy. The baseline degrades with time, while the gated model remains stable.

This matters because inference-time compute is only useful if extra computation remains coherent. A model that gets worse when allowed to continue is not thinking longer. It is dissolving.

Early exits improved

Because the modified model remains accurate across the trajectory, it also works better with early-exit rules. If we stop once confidence crosses a threshold, the gated model can reach that threshold quickly without paying as much in accuracy.

First Certainty
Figure 5. First crossing of the 0.8 certainty threshold. The modified model reaches usable confidence more cleanly.

This is the nice version of inference-time scaling: not “always spend more compute,” but “spend compute until the internal state has settled.”

The gate learned the story

Finally, I looked at rtr_t. If the whole setup was doing what I hoped, retention should start low and rise over time. That is what happened. The curve has the shape one would want: early openness to perception, then increasing internal closure.

Retention Panels
Figure 6. Retention over time. The model learns a clean transition from looking outward to dwelling inward.

Why I care about this

The point is not CIFAR-10. The point is that “thinking time” should be an object we can shape.

A lot of current scaling feels like increasing the size of the mouth: more tokens, longer traces, larger contexts. The CTM points toward a different axis. Give the model an internal temporal workspace, then ask how it should regulate contact with the world.

The Perceptual Gate is a small experiment in that direction. It is probably not the final mechanism. It may not even be the right mechanism. But it makes one thing visible: a model can learn when to stop looking and start dwelling.

That seems like a primitive worth taking seriously.