OpenAI unveils Jalapeño, its first in-house inference chip, and rates it ahead of the best hardware on the market
Jalapeño is a chip designed by OpenAI for inference (running an already-trained model so it produces its answers). On SemiAnalysis's independent InferenceX benchmark, it delivers both more tokens per user and more throughput per kilowatt than the best hardware available, according to the tests reported by TechCrunch and The Verge. Richard Ho, OpenAI's head of hardware, claims "the best of both worlds," lower latency and higher throughput at the same time. The announcement comes with a piece by chief financial officer Sarah Friar on "the full stack behind abundant intelligence": OpenAI wants to own the entire chain, from chips to product, to bring down the cost of each answer. Designing its own inference silicon is a bid to loosen its dependence on Nvidia and take back control of the bill weighing on its margins.

