FOUNDER FOCUSWe speak with AI founders and executives about AI, startups, and whatever else is on their minds. Nearly every leading chatbot you use today runs on the same piece of math. The transformer, published by Google researchers in 2017, does one thing: it predicts the next word. Do that well enough, billions of times over, and you get ChatGPT. Now, a Palo Alto AI lab is trying to build an alternative. Pathway calls it the post-transformer world. We spoke with co-founder and CEO Zuzanna Stamirowska to know more. The interview has been edited for brevity and clarity.  Why do you think AI needs to move beyond transformers? This is math made to predict sequences. In language, you're trying to guess the next word. And this math is limited. Right now, transformers have no memory after they are trained. You train a model, the weights are set. As you use it, you deal with context, but that doesn't impact the weights anymore. So what transformer models do is relive their Groundhog Day, or the Memento movie. They wake up every day with a fresh mind, without having internalized anything that happened to them. Acquiring new skills, acquiring judgment, is something you do through experience. Transformers fundamentally don't do that. If transformers are limited, what does a better approach to reasoning look like? Reasoning was bolted on to the transformer with a chain of thought, a process in which every step of thinking has to be verbalized in language. That constrains thinking to language. What we discovered at Pathway is that thinking doesn't happen in language. It happens in some sort of abstract space. Once you release reasoning from the constraints of language, you can do way more powerful things way faster. Sudoku is a good example. It's difficult to express in language sequentially, but it's obvious as you see it. Your grandmother can solve it over coffee, and yet it's deeply difficult for LLMs to do natively. They do it by finding tricks. This matters because compute is exploding. Anthropic would have had even larger growth if it had more compute. A lot of that compute is going toward reasoning, because the chain of thought requires so much of it. If you move to latent reasoning, you cut these costs dramatically. For us, solving Sudoku - both training the model and running it - costs so little it was unworthy of reporting. Can you explain the alternative architecture Pathway is building? We published a paper in October titled "The dragon hatchling: the missing link between transformer and the brain." It was the second most popular paper on Hugging Face in 2025. Everything is organized as neurons, small computational units connected through synapses, forming a network like a brain. Not everybody is connected to everybody. You have your neighbors, a bit like in a social network. When a signal comes in, the neuron that gets activated pushes it only to its neighbors. If they find it relevant, they fire up too. And if both neurons at the ends fire up, that connection becomes stronger. This forms a memory. So memory is stored directly on the synapses within the model. In transformers, you have fixed context windows. Here, you have access to the entire context all the time, but you only trigger the synapses relevant to your query. That gives us incredible efficiency. Does moving beyond transformers mean changing the hardware stack? That was the key part. Transformers sort of won the hardware lottery, and that's great, because it allowed scale. But we made this architecture work on any GPU. We mostly work on H200s. Will it work better one day on specialized chips? Yes, but you don't need any special chip. If we had to change the chip and wait for everything to align, that would be a difficult path. Where has the most interesting post-transformer research been coming from? Quite the opposite of what you'd expect. The US is the most transformer-religious place. There was a time when most of the papers I was reading were from China, because there was no religion about doing just the transformer, and it was far more innovative. The first interesting work on sparsity that we saw came out of China. The US was very focused on brute-force scaling for a long time. How close are you to bringing this to market? Right now, we are a lab. We invest in our moat. The more time we spend being dramatically differentiated, the bigger the impact. We have a partnership with AWS and Nvidia, and we're working with design partners on personalized AI and long-horizon reasoning. On cost, it's at least 10x less than open models for advanced reasoning. We should have public numbers soon. Strategically, it makes little sense to build yet another OpenAI or yet another Anthropic. A breakthrough is inevitable - the markets just don't see it yet. |