For years, the industry’s easiest path to faster chips was simple: crank up the clock speed. That worked great, until it didn’t. By the early 2000s, power consumption and heat had put a hard ceiling on performance. Moore’s Law, at least in its easy-gains form, was over.
Most of the industry scrambled for patches. Kunle Olukotun had already bet on something bigger: put multiple processors on one chip. That idea reshaped how modern processors get built. Decades later, it quietly became one of the engines behind the AI boom.
Betting Against the Bigger-Core Race
Back in 1996, everyone was building wider, more complex processors—more instruction issue, more speculative execution, more silicon dedicated to squeezing parallelism out of a single stream of instructions.
Olukotun, along with four Stanford colleagues, thought this was a dead end. Their pitch was simple: take one complicated processor, replace it with four simpler ones, keep the silicon area roughly the same. They ran the numbers assuming equal 500 MHz clocks and similar 430 mm² dies. The catch? For anything that couldn’t be parallelized, the old superscalar design still won—by about 30%. But for parallel and multiprogramming workloads, the four-core setup crushed it, beating the single-core design by 50% to 100%.
Olukotun wasn’t content to leave the idea in academia. He turned it into Stanford’s Hydra chip-multiprocessor project before founding Afara Websystems to bring the concept to market. Sun Microsystems bought Afara in 2002. That technology became the backbone of Niagara. By December 2005, Sun was shipping the UltraSPARC T1—eight cores, 32 simultaneous threads. According to the IEEE Computer Society, Niagara-derived chips eventually pulled in billions in revenue.
Here’s the thing: the bet wasn’t “multicore will be faster.” It was that future software would have enough parallelism baked in to make many simple cores beat one very complicated core. That’s a much riskier bet. It paid off anyway.
Design Around the Constraint, Not the Fashion
Find what actually limits performance: Don’t chase the metric everyone’s talking about. Find the resource that gets scarce as things scale.
“The inference problem is not really a compute problem.” That’s Olukotun, speaking on AI:AM in July 2026. His argument now: AI inference is bottlenecked by moving model parameters and KV-cache data around, not by raw FLOPS. SambaNova’s dataflow architecture is built around exactly that—maximizing memory-bandwidth use and overlapping communication with computation.
Don’t specialize past the market: Specialization should cut waste, not lock you into an assumption that might not hold in five years. “I’ve learned never to bet against the innovation capabilities of software people and algorithm people,” he’s said. SambaNova went with reconfigurable dataflow instead of hard-wiring a single algorithm into silicon. The chip can specialize, sure—but it can also adapt as models evolve.
Push complexity away from the user: Even great technology stalls if people have to babysit its internals. “We need to raise the level of abstraction for parallel programming,” Olukotun told The Register back in 2008. His Stanford Pervasive Parallelism Lab chased domain-specific languages and compilers that could find parallelism on their own—no manual thread-juggling required. Sun, AMD, Nvidia, IBM, HP, and Intel backed the lab with $6 million over three years.
When Afara Hit a Wall
Co-founder Les Kohn later remembered it plainly: Afara had closed one funding round, and then September 11 happened. Fundraising froze overnight. With fresh capital suddenly unavailable, Afara had little room to stay independent. Sun acquired the company, preserving the technology even as the startup disappeared.
There’s no solid evidence Olukotun ever pinned a specific lesson on that moment. But something did shift—the capital surrounding his work got a lot sturdier. PPL had six major tech companies behind it. SambaNova later drew SoftBank Vision Fund 2, BlackRock, Intel Capital, GV, Temasek, GIC, and others. In 2026 alone, SambaNova pulled in over $350 million in Series E funding, plus a planned strategic partnership with Intel.
The Network Behind the Architecture
DARPA funded the original 1996 research. Stanford gave him the platform to build on. Chipmakers funded PPL. Institutional investors funded SambaNova. One protégé, Hassan Chafi, did his PhD under Olukotun—working on transactional memory and the Delite DSL—before moving on to lead research at Oracle Labs.
Put it together, and you get something rare: academic research, chip talent, enterprise buyers, semiconductor relationships, and patient capital, all connected. That’s what it actually takes to commercialize an idea that might not look obvious for years.
Find the Constraint That Gets Worse With Scale
Before you optimize your product, do this instead: find the one resource whose cost—or inefficiency—grows fastest when usage jumps 10x. Then run one experiment aimed squarely at cutting that in half.
That’s the whole Olukotun pattern, really. Don’t polish an architecture whose bottleneck is already structural. Replace it.
