Skip to main content

Posts

Showing posts with the label deep learning research

The n_embd Confound: A Micro Scaling-Law Sweep on CW-Node

This is a write-up of a small, honest experiment: does splitting a transformer's feed-forward block into a shared "external" MLP and a private per-channel "internal" MLP (an architecture I'm calling CW-Node) buy any parameter efficiency over a plain dense baseline? Twelve runs, four parameter tiers, three architectures. Short answer: not yet, and the reason why turned out to be more interesting than the result itself. What CW-Node changes Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures is entirely inside that feed-forward block. A standard dense block runs one shared MLP across the full feature vector at every position — a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works. CW-Node...