On the Infinite Width and Depth Limits of Predictive Coding Networks

Achour EM

Modern AI systems like ChatGPT are built on neural networks—structures of interconnected artificial neurons arranged in layers, loosely inspired by the biological brain. Currently, these networks are trained using an algorithm called “backpropagation”. While highly effective, backpropagation requires distant neurons to communicate with one another, which not only doesn’t align with how the brain actually learns, but is also incredibly energy-inefficient. A alternative algorithm called “predictive coding” is much more brain-like because it only updates connections based on the activity of neighboring neurons.

However, a major question remains: can predictive coding scale up to match the performance of massive modern AI models? We answer this question by mathematically analysing what happens when predictive coding networks become incredibly wide (many neurons per layer) and deep (many layers).

We show that when a predictive coding network is designed to be much wider than it is deep—closely mirroring the actual structural proportions of the human brain—it updates its connections in the exact same way as backpropagation. This work suggests how the brain could learn effectively using only local updates, while contributing to the development of scalable, energy-efficient AI.

Scientific Abstract

Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of deep PC networks (PCNs) by leveraging some BP-inspired reparameterisations, but the scalability and theoretical basis of these methods remain unclear. To address this gap, we study the infinite width and depth limits of PCNs. For linear networks, we derive stable and ``non-lazy'' parameterisations when scaling both the model width and depth, revealing that the output of standard PCNs explodes with width during training. Moreover, under stable parameterisations, we show that the gradients computed by PC at activity equilibrium converge to the BP gradients for networks that are much wider than deep (depth/width → 0). Experiments show high gradient alignment between PC and BP at large width for different nonlinear models, including convolutional networks and transformers. Overall, this work constrains the parameterisations that are scalable with PC, while suggesting how BP could be implemented using only local updates in much wider than deep networks like the brain.

Similar content

Preprint
MRC CoRE in RND Output
Wang Y, Burgeno L, Cerpa JC, Manohar S, Bogacz R, Walton ME

Temporally distinct reward and action prediction error signals during value learning and habit formation

On the Infinite Width and Depth Limits of Predictive Coding Networks

Achour EM

Modern AI systems like ChatGPT are built on neural networks—structures of interconnected artificial neurons arranged in layers, loosely inspired by the biological brain. Currently, these networks are trained using an algorithm called “backpropagation”. While highly effective, backpropagation requires distant neurons to communicate with one another, which not only doesn’t align with how the brain actually learns, but is also incredibly energy-inefficient. A alternative algorithm called “predictive coding” is much more brain-like because it only updates connections based on the activity of neighboring neurons.

However, a major question remains: can predictive coding scale up to match the performance of massive modern AI models? We answer this question by mathematically analysing what happens when predictive coding networks become incredibly wide (many neurons per layer) and deep (many layers).

We show that when a predictive coding network is designed to be much wider than it is deep—closely mirroring the actual structural proportions of the human brain—it updates its connections in the exact same way as backpropagation. This work suggests how the brain could learn effectively using only local updates, while contributing to the development of scalable, energy-efficient AI.

Scientific Abstract

Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of deep PC networks (PCNs) by leveraging some BP-inspired reparameterisations, but the scalability and theoretical basis of these methods remain unclear. To address this gap, we study the infinite width and depth limits of PCNs. For linear networks, we derive stable and ``non-lazy'' parameterisations when scaling both the model width and depth, revealing that the output of standard PCNs explodes with width during training. Moreover, under stable parameterisations, we show that the gradients computed by PC at activity equilibrium converge to the BP gradients for networks that are much wider than deep (depth/width → 0). Experiments show high gradient alignment between PC and BP at large width for different nonlinear models, including convolutional networks and transformers. Overall, this work constrains the parameterisations that are scalable with PC, while suggesting how BP could be implemented using only local updates in much wider than deep networks like the brain.

Citation

2026, Proceedings of the 43rd International Conference on Machine Learning

Downloads

View PDF (3MB)

Similar content

Preprint
MRC CoRE in RND Output
Wang Y, Burgeno L, Cerpa JC, Manohar S, Bogacz R, Walton ME

Temporally distinct reward and action prediction error signals during value learning and habit formation