
Image: generated by METAL AI
Summary
- OpenAI CFO Sarah Friar said in a September 8 blog post that the company will deploy its in-house inference chip, Halapeno, before the end of the year.
- She said InferenceX testing showed the chip delivering 1.5 to 1.9 times the tokens-per-watt throughput of commercial systems, with 1.7 to 3.6 times lower latency.
- She framed this as part of a flywheel in which consumer and enterprise revenue reinforce each other, built on a base of 1 billion weekly active users and 2.5 million enterprise customers.
OpenAI CFO Sarah Friar wrote in a company blog post on September 8 that the company plans to put its in-house inference chip, "Halapeno," into production before the end of the year. Her broader argument is that revenue from consumers and enterprise customers reinforces itself, feeding back into compute and research investment — and that this cycle sits at the core of OpenAI's growth strategy. The chip performance figures disclosed for the first time in this post were presented as evidence backing that claim.
This isn't a new model announcement — it's a strategy essay filed under the company's corporate blog category. Friar has written repeatedly about compute strategy and financial structure since last August, and this post ends with a list of her earlier entries. In other words, this isn't a one-off announcement but another chapter in a narrative OpenAI keeps retelling to investors and partners.
To put that in context: OpenAI released a new model, GPT-6 Astra, on September 3, and two days later that model became the first to receive a "critical" cybersecurity rating under OpenAI's own classification system. OpenAI's research team had also already reported that its coding agent uses the equivalent of 3.1 days of agent labor per day of human labor. This post revisits both of those data points and adds the Halapeno chip to the mix to explain how OpenAI is managing revenue and cost together.
Start with the scale Friar cited: across ChatGPT, ChatGPT Work, and Codex, weekly active users have topped 1 billion, and enterprise customers now number 2.5 million. Internal research tracking individual users found that daily message volume six months after sign-up runs about 50% higher than in the first month, and the range of tasks users try roughly doubles over that period. Her explanation is that free, ad-supported access introduces people to what AI can do, and they then increase spending through subscriptions and usage-based plans — a pattern that underpins this growth.
Cost figures came alongside the growth numbers. GPT-5.6 Sol has been used to improve production serving software, cutting end-to-end serving costs by 20%, and further improvements raised token-generation efficiency by more than 15%. That means OpenAI can extract more output from the same compute — and those software-side savings dovetail with hardware-side savings, which brings us to the post's headline news: Halapeno.
Halapeno is OpenAI's first custom-built inference chip. Running three publicly available models through InferenceX testing, and normalizing for rated power, OpenAI said the chip delivered 1.5 to 1.9 times the peak tokens-per-watt throughput of commercial systems, with end-to-end latency 1.7 to 3.6 times lower. OpenAI said it plans to deploy the chip by the end of the year alongside accelerators from partners including NVIDIA and AMD — meaning it's building its own silicon without fully replacing existing suppliers, at least for now.
In the OpenAI blog post, Friar wrote that "better models unlock new work." She went on to describe a loop in which compute efficiency lets that work scale up, and the resulting revenue growth funds the next round of research and infrastructure. She added that OpenAI weighs each investment against the demand it will meet, how quickly it converts into productive capacity, and whether the returns justify the capital deployed — language that reads as a preemptive answer to market questions about the company's large data-center commitments.
The timing of this message is worth noting. Both OpenAI and Anthropic cut prices last month in response to competition from Chinese AI rivals, and OpenAI lowered GPT-5.6 Luna's token pricing by 80% for both input and output. Cutting prices while protecting profitability requires actually reducing serving costs, and the 20% cost cut and Halapeno's performance figures highlighted in this post amount to OpenAI's answer to that challenge. Corporate partnerships with OpenAI have also been growing — Samsung SDS, for instance, recently said it became the first company in Korea to join OpenAI's Daybreak program — and as OpenAI drives compute costs down further, the pricing competitiveness of these enterprise offerings is likely to shift as well.
The numbers in this post remain self-reported internal metrics for now. Whether Halapeno's real-world performance after deployment matches the InferenceX results, and how much of that 20% cost cut actually flows into lower prices, will become clear in next quarter's earnings and pricing decisions. In the end, the real test of this consumer-enterprise flywheel narrative isn't a flashy benchmark — it's whether OpenAI actually passes its cost savings on to customers.





Comments