Skip to main content

The n_embd Confound: A Micro Scaling-Law Sweep on CW-Node

This is a write-up of a small, honest experiment: does splitting a transformer's feed-forward block into a shared "external" MLP and a private per-channel "internal" MLP (an architecture I'm calling CW-Node) buy any parameter efficiency over a plain dense baseline? Twelve runs, four parameter tiers, three architectures. Short answer: not yet, and the reason why turned out to be more interesting than the result itself.

What CW-Node changes

Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures is entirely inside that feed-forward block.

A standard dense block runs one shared MLP across the full feature vector at every position — a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works.

CW-Node splits that block's parameter budget into two pools that run in sequence:

  • External pool — the same shared dense MLP as the baseline, with width w_ext and depth d_ext. It produces one output value per channel, mixing information across channels as usual.
  • Internal pool — each of those output channels then gets routed through its own private MLP, with its own weights, width w_int, and depth d_int. This MLP sees only that one channel's scalar value — no mixing across channels — and re-expands and contracts it before the block's final output.

The two variants tested, CW-Node 70/30 and CW-Node 30/70, name the approximate share of the block's feed-forward parameters given to the external pool versus the internal pool. 70/30 keeps most of the budget in shared, cross-channel mixing; 30/70 moves most of it into private, per-channel processing.

Setup

The corpus is a 5,000,000-token character-level split (500,000 held out for validation), vocabulary size 113. Every run trains for exactly 19,531 steps — one epoch over the training split at block size 64 and batch size 4 — using AdamW at a learning rate of 3e-3, on an Apple Silicon GPU via the mps backend.

Twelve runs cover three architectures at four total-parameter tiers: 500K, 1M, 3M, and 5M. The parameter solver holds n_embd = 16 fixed everywhere and adjusts w_ext / w_int per run to land each architecture within roughly ±1.5% of its tier's target.

Results

Lower validation loss is better. Delta is CW-Node minus dense at the same tier, in nats — a negative delta is a CW-Node win.

TierArchitectureParams (actual)Val lossΔ vs denseWall time
500KDense502,2722.6452—201s
500KCW-Node 70/30501,7542.6052-0.0400306s
500KCW-Node 30/70494,5542.6541+0.0089283s
1MDense1,003,1362.5470—256s
1MCW-Node 70/30999,4822.6805+0.1335443s
1MCW-Node 30/70997,0342.6020+0.0550516s
3MDense3,007,0762.5430—388s
3MCW-Node 70/302,996,2072.7756+0.2326507s
3MCW-Node 30/702,984,2762.5918+0.0488733s
5MDense5,007,1192.5800—489s
5MCW-Node 70/304,992,7402.5682-0.0118686s
5MCW-Node 30/704,983,8122.7140+0.1340906s

Dense wins six of eight matched comparisons. Both CW-Node wins are narrow and sit at opposite ends of the parameter range: 70/30 beats dense by 0.0400 nats at 500K and 0.0118 nats at 5M, with a loss of up to 0.2326 nats in between at 3M. 30/70 loses at every tier, by margins from 0.0089 to 0.1340 nats. Neither variant shows a consistent trend with parameter count in either direction.

Wall-clock cost

Parameter count isn't the only budget that matters. CW-Node's internal pool runs through a custom chunked autograd function rather than a single batched matrix multiply, and it costs real time on every run in this sweep — overhead relative to dense ranges from 1.30× to 2.02×, averaging roughly 1.6×, with 30/70 consistently the slower of the two variants. Some of this is an implementation cost rather than an architectural one: the internal pool's per-channel MLPs are computed with a hand-written chunked forward/backward pass built to work around an MPS compiler deadlock, not a kernel tuned for throughput. Even so, any efficiency claim for CW-Node has to clear this bar too, not just the parameter-count bar above.

The n_embd confound

The comparison above is fair on parameter count and unfair on the dimension that likely matters more for this architecture. The internal pool's per-channel MLP has an input of exactly one scalar and an output of one scalar. Its capacity to do anything beyond a smooth reshaping of that single number depends on the intermediate width w_int, which the solver has to keep small because n_embd is locked at 16 across every run in this sweep.

A dense baseline saturates a 16-wide embedding easily — its shared matrix multiply mixes all 16 channels directly. CW-Node's internal pool, forced into the same 16-wide space, spends part of its budget on private per-channel machinery that a starved embedding gives little to work with. That reading is consistent with the data, but it's a plausible explanation given the constraint, not a claim demonstrated by a controlled comparison — this sweep never varies n_embd, so it can't isolate that variable from parameter allocation directly.

The concrete next test is to let the solver vary n_embd per architecture under a fixed total-parameter budget instead of pinning it. At a 3M-parameter budget, for example, a dense model might land near n_embd = 128 while a CW-Node model might land near n_embd = 384. If CW-Node still loses once it's allowed a wider embedding, that's real evidence against the architecture. If it wins, that shows the internal-routing pathway has a genuine efficiency advantage that this sweep's fixed embedding simply couldn't expose.

Limitations

  • Single seed per cell. All twelve runs are one training run each — no repeated seeds, no error bars. Differences under roughly 0.03–0.05 nats, including both CW-Node "wins," are within a range seed variance alone could plausibly account for.
  • Character-level, small vocabulary. A 113-symbol vocabulary and a 5M-token corpus are a narrow language-modeling regime. Whether the same trade-off holds at subword-tokenized scale or on larger corpora is untested here.
  • Wall-clock reflects one implementation. The overhead numbers above are specific to the current chunked-autograd internal pool, not necessarily to internal-routing as an idea in general.

Where this goes next

It's early, and honestly not promising yet — dense is winning most of the comparisons. The next sweep removes the fixed-n_embd constraint: the parameter solver varies embedding width per architecture under a fixed total-parameter budget, so each architecture gets the embedding width its own scaling behavior calls for. That's the direct test of whether the internal-routing pathway is disadvantaged by design or was simply denied the representational room to show what it does.

Full data, charts, and reproduction instructions are on the project page: mohamedhossammohamed.github.io/cw-node-research. Code and raw results are in the repository: github.com/mohamedhossammohamed/cw-node-research.

Questions, pushback, or ideas for the next sweep — find me on X: @MohamedHz72007.

Comments

Popular posts from this blog

Trying to Speed Up a Weird Neural Net Idea on My Mac's GP

I've been playing around with a transformer variant where, instead of one shared feed-forward layer, every single neuron gets its own tiny private network. It's a fun idea, but the first version I wrote ran painfully slow on my Mac, and this post is basically my notes on figuring out why, and what I tried to fix it. The Setup The model has 3 layers and around 2,574 "neurons" at full size (I mostly tested a shrunk-down version with way fewer, just to iterate faster). Each neuron has its own small 2-layer MLP with its own weights — nothing is shared between neurons at this stage. Each layer does roughly: Normal causal self-attention A regular shared MLP that mixes information across neurons The weird part: each of the ~2,574 neurons runs through its own tiny 2-layer network, completely independently for each neuron n: h = GELU( x[n] * in_w[n] + in_b[n] ) h = GELU( LN( h @ hw0[n] + hb0[n] ) * lw0[n] + lb0[n] ) + h h = GELU( LN( h @ hw1[n] + hb1[...

Outside the X bubble of software developers and tokenmaxers, how are non-technical professionals actually using AI in production to run real, revenue-generating businesses?

If you scroll tech Twitter, the discourse is dominated by "tokenmaxing" — people flexing multi-million token usage, chaining complex autonomous agent loops, and obsessing over API token throughput. But far away from that hype, how does AI function in a real enterprise setting where revenue, accuracy, and client relationships are on the line? I recently had a deep conversation with a medical professional who offers a grounded perspective missing from mainstream tech discussions. He is a co-founder of a medical communications agency based in Jeddah, working face-to-face with major pharmaceutical companies and healthcare professionals. His agency bridges the gap between medical science and advertising — building research-backed, highly regulated campaigns. Having worked in the industry for over a decade — spanning 9-to-5 corporate roles, freelancing, and now running an agency — his experience provides a clear benchmark for where AI provides genuine leverage, where it expos...

Shovels, Chatbots, and the AI Bubble Nobody Else Is In

I talked to a pharma executive last week. Big company, UK based, he runs things from what I think is the Middle East control room in Jeddah. Senior guy, more than a decade in the industry. I asked him the obvious question, how is he using AI day to day. He does not, not really. He uses Copilot to proofread emails and juggle the occasional idea, because Copilot is the only thing his work laptop will let him touch. He told me he thought about running a second laptop off the company firewall so he could actually explore what is out there, but keeping two work laptops was too much of a hassle. So he stayed on one laptop, inside the firewall, and stayed a chatbot user. That is not a story about one lazy executive. That is a story about an entire class of sectors, heavy IP, heavy security, heavy regulation, where the friction of adopting anything beyond a sanctioned chatbot is high enough that the frontier simply does not reach them yet. Pharma, medicine, anywhere production touches actu...