OpenAI’s chip chief tells More Than Moore he doesn’t fully trust AI design
OpenAI’s own AI models changed parts of its Jalapeño chip, and its engineers could not always say why. In one example, they saved more than 13% of die area, Richard Ho told Ian Cutress of More Than Moore. Ho leads OpenAI’s hardware team, and the interview was published on 30 September. “Well, I still don’t […] This story continues at The Next Web
OpenAI’s own AI models changed parts of its Jalapeño chip, and its engineers could not always say why. In one example, they saved more than 13% of die area, Richard Ho told Ian Cutress of More Than Moore . Ho leads OpenAI’s hardware team, and the interview was published on 30 September.
“Well, I still don’t 100% trust them,” Ho told More Than Moore.
OpenAI revealed its own AI chip in June. It built the chip with Broadcom to run its models, not to train them. At Hot Chips in August, it showed benchmark results against Nvidia’s GB300. The interview is about how the chip was made. Ho joined OpenAI in 2023. He was one of the first engineers on Google’s TPU programme. Before that, he built the Anton supercomputers at D.E. Shaw Research.
The team moved fast, Ho said. Its early estimates of how much logic would fit on the chip proved a little off. Rather than give up performance, it handed the problem to its models. At the time, he said, they were raw internal models, not fine-tuned for chip work. They were only slightly ahead of what the public could use.
Much of the code the models worked on was written in XLS, a tool that turns high-level code into chip logic. Chris Leary, now on Ho’s team, released it as open source while at Google. It reads more like Rust than Verilog, the language most chips are written in. The models handled software better than Verilog, Ho said, so XLS let them reason about the design.
Some of the models’ changes to the source code came without a clear reason. The team was running too fast to stop and work out why they helped, Ho said.
“We didn’t quite understand conceptually what it was doing,” he said.
So every change still went through the full validation flow, and standard design tools still sign off the chip. A chip can have hundreds of millions of gates. A model that is right 99.99% of the time is not accurate enough for that, he said. He called it a risk worth taking, as long as the team stays cautious. He would not cut staff because of the models, he added. He would use them to do more, faster.
Ho also defended the chip’s most unusual choice. Much of the industry now splits inference across separate hardware for each stage: prefill, which reads the prompt, and decode, which writes the answer.
OpenAI built one part for all of it. Locking data centre capacity into a fixed ratio is risky when the mix of work keeps changing, Ho said. He conceded the design might carry a cost that only shows up in real deployment, and that very long contexts could hit limits.
The chip links banks of high-bandwidth memory directly to its cores, so data moves less. Ho expects rivals, GPU makers included, to copy the idea. OpenAI chose to publish it anyway, even though it is among Nvidia’s biggest customers .
“Because we still need GPUs, and we want GPUs to get better,” Ho said.
In his view, the industry’s real limit is not power, memory or packaging. It is that executives and investors, burned by past booms, cannot see how big AI will get, so their bets are too small. The second version of Jalapeño is now in the lab for qualification, as OpenAI moves toward volume production.