Skip to main content

Posts

Showing posts with the label LLM

Shovels, Chatbots, and the AI Bubble Nobody Else Is In

I talked to a pharma executive last week. Big company, UK based, he runs things from what I think is the Middle East control room in Jeddah. Senior guy, more than a decade in the industry. I asked him the obvious question, how is he using AI day to day. He does not, not really. He uses Copilot to proofread emails and juggle the occasional idea, because Copilot is the only thing his work laptop will let him touch. He told me he thought about running a second laptop off the company firewall so he could actually explore what is out there, but keeping two work laptops was too much of a hassle. So he stayed on one laptop, inside the firewall, and stayed a chatbot user. That is not a story about one lazy executive. That is a story about an entire class of sectors, heavy IP, heavy security, heavy regulation, where the friction of adopting anything beyond a sanctioned chatbot is high enough that the frontier simply does not reach them yet. Pharma, medicine, anywhere production touches actu...

Trying to Speed Up a Weird Neural Net Idea on My Mac's GP

I've been playing around with a transformer variant where, instead of one shared feed-forward layer, every single neuron gets its own tiny private network. It's a fun idea, but the first version I wrote ran painfully slow on my Mac, and this post is basically my notes on figuring out why, and what I tried to fix it. The Setup The model has 3 layers and around 2,574 "neurons" at full size (I mostly tested a shrunk-down version with way fewer, just to iterate faster). Each neuron has its own small 2-layer MLP with its own weights — nothing is shared between neurons at this stage. Each layer does roughly: Normal causal self-attention A regular shared MLP that mixes information across neurons The weird part: each of the ~2,574 neurons runs through its own tiny 2-layer network, completely independently for each neuron n: h = GELU( x[n] * in_w[n] + in_b[n] ) h = GELU( LN( h @ hw0[n] + hb0[n] ) * lw0[n] + lb0[n] ) + h h = GELU( LN( h @ hw1[n] + hb1[...