I've been playing around with a transformer variant where, instead of one shared feed-forward layer, every single neuron gets its own tiny private network. It's a fun idea, but the first version I wrote ran painfully slow on my Mac, and this post is basically my notes on figuring out why, and what I tried to fix it. The Setup The model has 3 layers and around 2,574 "neurons" at full size (I mostly tested a shrunk-down version with way fewer, just to iterate faster). Each neuron has its own small 2-layer MLP with its own weights — nothing is shared between neurons at this stage. Each layer does roughly: Normal causal self-attention A regular shared MLP that mixes information across neurons The weird part: each of the ~2,574 neurons runs through its own tiny 2-layer network, completely independently for each neuron n: h = GELU( x[n] * in_w[n] + in_b[n] ) h = GELU( LN( h @ hw0[n] + hb0[n] ) * lw0[n] + lb0[n] ) + h h = GELU( LN( h @ hw1[n] + hb1[...