This is a write-up of a small, honest experiment: does splitting a transformer's feed-forward block into a shared "external" MLP and a private per-channel "internal" MLP (an architecture I'm calling CW-Node) buy any parameter efficiency over a plain dense baseline? Twelve runs, four parameter tiers, three architectures. Short answer: not yet, and the reason why turned out to be more interesting than the result itself.
What CW-Node changes
Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures is entirely inside that feed-forward block.
A standard dense block runs one shared MLP across the full feature vector at every position — a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works.
CW-Node splits that block's parameter budget into two pools that run in sequence:
- External pool — the same shared dense MLP as the baseline, with width
w_extand depthd_ext. It produces one output value per channel, mixing information across channels as usual. - Internal pool — each of those output channels then gets routed through its own private MLP, with its own weights, width
w_int, and depthd_int. This MLP sees only that one channel's scalar value — no mixing across channels — and re-expands and contracts it before the block's final output.
The two variants tested, CW-Node 70/30 and CW-Node 30/70, name the approximate share of the block's feed-forward parameters given to the external pool versus the internal pool. 70/30 keeps most of the budget in shared, cross-channel mixing; 30/70 moves most of it into private, per-channel processing.
Setup
The corpus is a 5,000,000-token character-level split (500,000 held out for validation), vocabulary size 113. Every run trains for exactly 19,531 steps — one epoch over the training split at block size 64 and batch size 4 — using AdamW at a learning rate of 3e-3, on an Apple Silicon GPU via the mps backend.
Twelve runs cover three architectures at four total-parameter tiers: 500K, 1M, 3M, and 5M. The parameter solver holds n_embd = 16 fixed everywhere and adjusts w_ext / w_int per run to land each architecture within roughly ±1.5% of its tier's target.
Results
Lower validation loss is better. Delta is CW-Node minus dense at the same tier, in nats — a negative delta is a CW-Node win.
| Tier | Architecture | Params (actual) | Val loss | Δ vs dense | Wall time |
|---|---|---|---|---|---|
| 500K | Dense | 502,272 | 2.6452 | — | 201s |
| 500K | CW-Node 70/30 | 501,754 | 2.6052 | -0.0400 | 306s |
| 500K | CW-Node 30/70 | 494,554 | 2.6541 | +0.0089 | 283s |
| 1M | Dense | 1,003,136 | 2.5470 | — | 256s |
| 1M | CW-Node 70/30 | 999,482 | 2.6805 | +0.1335 | 443s |
| 1M | CW-Node 30/70 | 997,034 | 2.6020 | +0.0550 | 516s |
| 3M | Dense | 3,007,076 | 2.5430 | — | 388s |
| 3M | CW-Node 70/30 | 2,996,207 | 2.7756 | +0.2326 | 507s |
| 3M | CW-Node 30/70 | 2,984,276 | 2.5918 | +0.0488 | 733s |
| 5M | Dense | 5,007,119 | 2.5800 | — | 489s |
| 5M | CW-Node 70/30 | 4,992,740 | 2.5682 | -0.0118 | 686s |
| 5M | CW-Node 30/70 | 4,983,812 | 2.7140 | +0.1340 | 906s |
Dense wins six of eight matched comparisons. Both CW-Node wins are narrow and sit at opposite ends of the parameter range: 70/30 beats dense by 0.0400 nats at 500K and 0.0118 nats at 5M, with a loss of up to 0.2326 nats in between at 3M. 30/70 loses at every tier, by margins from 0.0089 to 0.1340 nats. Neither variant shows a consistent trend with parameter count in either direction.
Wall-clock cost
Parameter count isn't the only budget that matters. CW-Node's internal pool runs through a custom chunked autograd function rather than a single batched matrix multiply, and it costs real time on every run in this sweep — overhead relative to dense ranges from 1.30× to 2.02×, averaging roughly 1.6×, with 30/70 consistently the slower of the two variants. Some of this is an implementation cost rather than an architectural one: the internal pool's per-channel MLPs are computed with a hand-written chunked forward/backward pass built to work around an MPS compiler deadlock, not a kernel tuned for throughput. Even so, any efficiency claim for CW-Node has to clear this bar too, not just the parameter-count bar above.
The n_embd confound
The comparison above is fair on parameter count and unfair on the dimension that likely matters more for this architecture. The internal pool's per-channel MLP has an input of exactly one scalar and an output of one scalar. Its capacity to do anything beyond a smooth reshaping of that single number depends on the intermediate width w_int, which the solver has to keep small because n_embd is locked at 16 across every run in this sweep.
A dense baseline saturates a 16-wide embedding easily — its shared matrix multiply mixes all 16 channels directly. CW-Node's internal pool, forced into the same 16-wide space, spends part of its budget on private per-channel machinery that a starved embedding gives little to work with. That reading is consistent with the data, but it's a plausible explanation given the constraint, not a claim demonstrated by a controlled comparison — this sweep never varies n_embd, so it can't isolate that variable from parameter allocation directly.
The concrete next test is to let the solver vary n_embd per architecture under a fixed total-parameter budget instead of pinning it. At a 3M-parameter budget, for example, a dense model might land near n_embd = 128 while a CW-Node model might land near n_embd = 384. If CW-Node still loses once it's allowed a wider embedding, that's real evidence against the architecture. If it wins, that shows the internal-routing pathway has a genuine efficiency advantage that this sweep's fixed embedding simply couldn't expose.
Limitations
- Single seed per cell. All twelve runs are one training run each — no repeated seeds, no error bars. Differences under roughly 0.03–0.05 nats, including both CW-Node "wins," are within a range seed variance alone could plausibly account for.
- Character-level, small vocabulary. A 113-symbol vocabulary and a 5M-token corpus are a narrow language-modeling regime. Whether the same trade-off holds at subword-tokenized scale or on larger corpora is untested here.
- Wall-clock reflects one implementation. The overhead numbers above are specific to the current chunked-autograd internal pool, not necessarily to internal-routing as an idea in general.
Where this goes next
It's early, and honestly not promising yet — dense is winning most of the comparisons. The next sweep removes the fixed-n_embd constraint: the parameter solver varies embedding width per architecture under a fixed total-parameter budget, so each architecture gets the embedding width its own scaling behavior calls for. That's the direct test of whether the internal-routing pathway is disadvantaged by design or was simply denied the representational room to show what it does.
Full data, charts, and reproduction instructions are on the project page: mohamedhossammohamed.github.io/cw-node-research. Code and raw results are in the repository: github.com/mohamedhossammohamed/cw-node-research.
Questions, pushback, or ideas for the next sweep — find me on X: @MohamedHz72007.
Comments
Post a Comment