TL;DR. A transformer’s opaque serial depth scales with its number of layers, $\sim L$. Any longer serial computation has to go through the tokens it writes. For illustration, I briefly discuss a subtly different, hypothetical kind of attention that reads the layer it is writing. That variant would have greater opaque serial depth, $\sim L+T$.

On the 80,000 Hours podcast, Rohin Shah describes today’s transformers as wide but shallow. A single forward pass does an enormous amount of work in parallel, but only a handful of steps in sequence. Reasoning that needs a longer chain has to spill into the chain of thought, where it can be read. That is a big part of why he is cautiously optimistic about chain-of-thought monitoring. In a paper with Jonah Brown-Cohen and David Lindner, he makes the intuition precise as opaque serial depth: the longest computation a model can do without interpretable intermediate steps.

Definition. Depth is circuit depth, the longest path through a circuit of binary associative ops and piecewise-analytic scalar functions. It is minimised over polynomial-size circuits, and in practice upper-bounded. A sum over $n$ inputs costs $\log_2 n$. Tokens (input, output, chain of thought) count as interpretable, and only paths between them count.

Residual stream grid, layers up, positions across. (a) Standard attention: every edge goes up a layer, so the longest opaque path is bounded by L. (b) Same-layer attention lets the path step right too, zig-zagging to about L+T.

Transformer (a). An opaque path through the residual stream $\boldsymbol h^\ell_t$ can only go up or right. Every attention edge also climbs a layer, so a path has at most $L$ steps, each costing $O(\log T + \log D)$:

$$\text{depth} = O\big(L(\log T + \log D)\big).$$

So a longer context barely helps. The only edge back down to layer 0 runs through a sampled token. Hence any computation in (a) that needs more serial depth than one forward pass provides ($\sim L$ layers) must be exposed in the emitted tokens (though it might be obfuscated). Those tokens may be compressed or steganographic, but they travel through a channel we can see.

Recurrence (b). Let position $t$ read position $t-1$ within a layer. The paper uses RNN blocks; panel (b) uses same-layer attention. Now a path takes $L$ steps up and $T$ steps right:

$$\text{depth} = O\big((L+T)\log D\big).$$

In code, take causal linear attention (no softmax), $\mathrm{Att}(\boldsymbol h)_t = \sum_{t'\le t} (\boldsymbol q_t \cdot \boldsymbol k_{t'})\, \boldsymbol v_{t'}$:

T, D = 8, 16                               # positions, width
W_q, W_k, W_v = (np.random.randn(D, D) / D**0.5 for _ in range(3))  # each (D, D)

def attn(h):                               # h: (t, D) -> (t, D)
    q, k, v = h @ W_q, h @ W_k, h @ W_v    # each (t, D)
    return np.tril(q @ k.T) @ v            # (t, t) causal scores @ (t, D) -> (t, D)

h0 = np.random.randn(T, D)                 # (T, D) token embeddings

The only difference between the two is where a layer reads from:

(a) attention only

h1 = attn(h0)  # (T, D) <- finished h0
h2 = attn(h1)  # (T, D) <- finished h1
h3 = attn(h2)  # (T, D), all t parallel
# depth ~ L

(b) + horizontal attention

h1 = attn(h0)                   # (T, D)
for t in range(T):              # serial in t
    h1[t] = attn(h1[:t+1])[-1]  # (t+1, D) -> (D,)
# repeat per layer; depth ~ L + T

That is linear in context, so long reasoning can stay hidden in the activations. This is why recurrence matters for safety: it would take away the guarantee that deep reasoning shows up in the output.


arXiv:2603.09786 · code · figure PDF / TikZ. The panel labels $\sim L$ and $\sim L+T$ leave out the log factors.