Skip to main content

Posts

Showing posts with the label Projects

The n_embd Confound: A Micro Scaling-Law Sweep on CW-Node

This is a write-up of a small, honest experiment: does splitting a transformer's feed-forward block into a shared "external" MLP and a private per-channel "internal" MLP (an architecture I'm calling CW-Node) buy any parameter efficiency over a plain dense baseline? Twelve runs, four parameter tiers, three architectures. Short answer: not yet, and the reason why turned out to be more interesting than the result itself. What CW-Node changes Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures is entirely inside that feed-forward block. A standard dense block runs one shared MLP across the full feature vector at every position — a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works. CW-Node...

Trying to Speed Up a Weird Neural Net Idea on My Mac's GP

I've been playing around with a transformer variant where, instead of one shared feed-forward layer, every single neuron gets its own tiny private network. It's a fun idea, but the first version I wrote ran painfully slow on my Mac, and this post is basically my notes on figuring out why, and what I tried to fix it. The Setup The model has 3 layers and around 2,574 "neurons" at full size (I mostly tested a shrunk-down version with way fewer, just to iterate faster). Each neuron has its own small 2-layer MLP with its own weights — nothing is shared between neurons at this stage. Each layer does roughly: Normal causal self-attention A regular shared MLP that mixes information across neurons The weird part: each of the ~2,574 neurons runs through its own tiny 2-layer network, completely independently for each neuron n: h = GELU( x[n] * in_w[n] + in_b[n] ) h = GELU( LN( h @ hw0[n] + hb0[n] ) * lw0[n] + lb0[n] ) + h h = GELU( LN( h @ hw1[n] + hb1[...