Skip to main content

Posts

Showing posts from July, 2026

Outside the X bubble of software developers and tokenmaxers, how are non-technical professionals actually using AI in production to run real, revenue-generating businesses?

If you scroll tech Twitter, the discourse is dominated by "tokenmaxing" — people flexing multi-million token usage, chaining complex autonomous agent loops, and obsessing over API token throughput. But far away from that hype, how does AI function in a real enterprise setting where revenue, accuracy, and client relationships are on the line? I recently had a deep conversation with a medical professional who offers a grounded perspective missing from mainstream tech discussions. He is a co-founder of a medical communications agency based in Jeddah, working face-to-face with major pharmaceutical companies and healthcare professionals. His agency bridges the gap between medical science and advertising — building research-backed, highly regulated campaigns. Having worked in the industry for over a decade — spanning 9-to-5 corporate roles, freelancing, and now running an agency — his experience provides a clear benchmark for where AI provides genuine leverage, where it expos...

The n_embd Confound: A Micro Scaling-Law Sweep on CW-Node

This is a write-up of a small, honest experiment: does splitting a transformer's feed-forward block into a shared "external" MLP and a private per-channel "internal" MLP (an architecture I'm calling CW-Node) buy any parameter efficiency over a plain dense baseline? Twelve runs, four parameter tiers, three architectures. Short answer: not yet, and the reason why turned out to be more interesting than the result itself. What CW-Node changes Every architecture in this sweep shares the same skeleton: causal single-head self-attention with a zero-initialized output projection, followed by a feed-forward block, stacked two or three times over a 16-dimensional embedding. The difference between the three architectures is entirely inside that feed-forward block. A standard dense block runs one shared MLP across the full feature vector at every position — a matrix multiply that mixes all channels together, the same way a Transformer FFN ordinarily works. CW-Node...

Trying to Speed Up a Weird Neural Net Idea on My Mac's GP

I've been playing around with a transformer variant where, instead of one shared feed-forward layer, every single neuron gets its own tiny private network. It's a fun idea, but the first version I wrote ran painfully slow on my Mac, and this post is basically my notes on figuring out why, and what I tried to fix it. The Setup The model has 3 layers and around 2,574 "neurons" at full size (I mostly tested a shrunk-down version with way fewer, just to iterate faster). Each neuron has its own small 2-layer MLP with its own weights — nothing is shared between neurons at this stage. Each layer does roughly: Normal causal self-attention A regular shared MLP that mixes information across neurons The weird part: each of the ~2,574 neurons runs through its own tiny 2-layer network, completely independently for each neuron n: h = GELU( x[n] * in_w[n] + in_b[n] ) h = GELU( LN( h @ hw0[n] + hb0[n] ) * lw0[n] + lb0[n] ) + h h = GELU( LN( h @ hw1[n] + hb1[...