LUA VISION

Maximum capacity is not maximum efficiency.

Capacity-Optimal Codes Are Not Energy-Optimal: The Synaptic Efficiency Coefficient as an Architectural Prior for Language Models. Paulo Câmara, LUA Vision, São Paulo, Brazil.

Today's AI gets better by getting bigger.

Every step in size costs more energy, more chips and more data centres. This work argues for another axis, one the brain has used for a long time: not storing more, but choosing better what to keep.

That matters to whoever pays the AI bill, to anyone who must run AI on their own servers, and to a country like Brazil, which will not win the race on size.

Two different optima.

The code with the largest capacity is not the most efficient once energy enters the objective. Optimising one does not bring you closer to the other.

Today's AI optimises the first and leaves the second for after training. By then the architecture has already been chosen under the wrong objective.

Far from the floor.

The physical minimum for erasing a bit is 3 × 10⁻²¹ J. A chemical synapse spends about 10⁴ ATP per bit, close to 10⁻¹⁵ J. Graded signals and spikes spend 10⁶ to 10⁷ ATP per bit, between 10⁻¹³ and 10⁻¹² J (Laughlin et al., 1998).

The best system nature has produced runs five to eight orders of magnitude above the limit. The design space is not exhausted.

The cost that did not fall.

FlashAttention cut the memory of attention from O(n²) to O(n). The operation count stays quadratic. It made the cost tolerable, not smaller.

SEC = α · β · γ

Pruning, potentiation and consolidation, in sequence, over nested synaptic populations. Each stage acts on what the previous one left.

A product, not a sum.

If one stage vanishes, the outcome must vanish. A sum does not respect that.

With a twenty per cent gain in each term, the product gives 1.728 and the sum 1.6.

Same substrate, different filter.

For α, dendritic spine density in BA21 (Tang et al., 2014): from childhood to adolescence the control group loses 45%, the ASD group 16%.

For β and γ, the valproate model (Markram and Markram, 2010): doubled LTP and fear memory resistant to extinction. Atypical development separates stages that move together in the typical case.

What would bring the coefficient down.

I propose measuring accuracy per joule. If the link between each term and the architectural decision does not survive that measure, the coefficient falls. The paper states these conditions.

How to cite.

Câmara, P. (2026). Capacity-Optimal Codes Are Not Energy-Optimal: The Synaptic Efficiency Coefficient as an Architectural Prior for Language Models (Version 1.1). LUA Vision Tecnologia. https://doi.org/10.5281/zenodo.22881788

@misc{camara2026sec,
  author    = {C{\^a}mara, Paulo},
  title     = {Capacity-Optimal Codes Are Not Energy-Optimal: The Synaptic
               Efficiency Coefficient as an Architectural Prior for
               Language Models},
  year      = {2026},
  version   = {1.1},
  publisher = {LUA Vision Tecnologia},
  doi       = {10.5281/zenodo.22881788},
  url       = {https://doi.org/10.5281/zenodo.22881788}
}