
Image: OpenAI
Summary
- OpenAI announced eight category winners of its Build Week challenge on August 25.
- Nearly 47,000 participants from 186 countries submitted more than 8,000 projects, competing for a total of $100,000 in prizes.
- Two of the four first-place teams were a veterinarian and a cardiologist with no coding background.
OpenAI announced the winners of its "Build Week" challenge on August 25 — the largest hackathon in the company's history, drawing nearly 47,000 participants from 186 countries.
There was one rule: build something real using Codex and GPT-5.6 and bring it back working. Over eight days, more than 8,000 projects were submitted, alongside seven online events and 60 in-person community meetups.
$100,000 in prizes, judged on four criteria
The challenge opened on July 13 and closed at 5 p.m. Pacific on July 21. A submission outage pushed the deadline back an hour, and the announcement itself slipped to August 25 after the volume of entries exceeded expectations.
Total prize money came to $100,000. Each category awarded $15,000 for first place and $10,000 for second, with first-place teams also receiving DevDay passes, a meeting with the Codex team, and a year of ChatGPT Pro. Judges scored entries on four criteria: technical execution, design and usability, potential impact, and quality of the idea.
There were four categories: Apps for Your Life, Work & Productivity, Developer Tools, and Education. Each category produced one first-place winner, one runner-up, and three finalists.
Apps for Your Life — giving voice to what couldn't be said
First place, Second Voice, was built by Ravitez Dondeti, who had a relative with a speech disability. As a child, Dondeti communicated with them through gestures, and only after they passed away did he realize how much had gone unasked. He built the project alone over the eight days of Build Week.
The app takes fragmented speech from someone with dysarthria or motor limitations, overlays it with the user's own phrasebook and the context of the moment, and surfaces just a handful of plausible sentences. Only after the user selects or edits one does the app speak it aloud. The person always keeps final control over what actually gets said.
Dondeti expected sentence reconstruction to be the hardest part. What actually held him up were the smallest decisions — how many options to show, how fast to surface them, how many taps confirmation should take.
"The smallest UX decisions carried the most weight. An app can be technically functional and still be useless if people can't actually use it."
Second place, AirBridge for Windows, grew out of a simple frustration: there was no way to use AirPlay speakers around the house from a Windows PC. Adam Tarantino had tried solving this once before with pre-2024 OpenAI models but never got it to a shippable product. This time, he went back at it.
The tool streams Windows audio to HomePods, Apple TV, and Macs — no virtual audio cables, no temp files. It supports simultaneous playback across multiple speakers, room-by-room latency correction, and a browser extension that delays video instead of audio to keep lip-sync intact. GPT-5.6 handles voice control of the system, a local policy layer decides which actions are allowed, and results get verified against the actual hardware.
"I've been building software for over ten years, but this was my first hackathon. You can't win if you don't show up."
Three finalists rounded out the category. Bander lets family members preview what an AI assistant plans to do before it touches an account. O2 by Agent9 pulls air quality, wildfire data, alerts, and forecasts into a single view. SayAhead helps deaf and hard-of-hearing users read, lead, and control phone calls.
Work & Productivity — when three vets become one
This year, Pauli VetCare went from three veterinarians down to one. The phone kept ringing with the same volume of calls, but there was no longer anyone to review every case.
First place, veTriage, was built by Erin Downes, the 61-year-old veterinarian who runs that clinic. She had no coding background. The tool helps front-desk staff ask the right intake questions, catch red-flag symptoms, and route cases to the right level of review — but it deliberately never lets non-clinical staff make medical judgment calls. The core insight: a full schedule doesn't make a patient less urgent, so clinical need and clinic capacity have to be handled as separate problems.
Codex turned Downes's own clinical judgment into a working workflow, and GPT-5.6 assists with intake conversations without making medical decisions itself. Downes's team is now piloting it in the actual clinic.
"At 61, what I want people to know is that it's not too late to become a builder if you have deep domain knowledge."
Second place, Pulse, is a research prototype built by Mohamed Mostafa Abu Taleb, a cardiologist in Cairo, between hospital shifts. During cardiac arrest resuscitation, the team lead has to simultaneously track rhythm and defibrillation, medication timing, CPR cycles, and interruptions in chest compressions — all while people are shouting information from every direction.
Pulse listens to the room and maintains a running clinical picture, built to handle speech that mixes Egyptian Arabic and English. GPT-5.6 interprets the chaotic speech, deterministic and auditable code tracks the resuscitation protocol itself, and the system asks the team to confirm anything it's uncertain about.
"No team, no lab, no funding — just built in Cairo, between hospital shifts."
Three finalists rounded out this category too. LabSpace AI turns a lab and everything in it into a searchable spatial map. OpenCounsel converts broken legal drafts into source-verified, filing-ready documents. Tomok stitches together schedules, specs, and reports so infrastructure teams can trace exactly where delays originated.
Developer Tools — sketching sound like a wireframe
Spatial audio usually only gets tested after a scene is fully built in a game engine — meaning there's no way to explore "how should this space sound?" early in the design process.
First place, Echo Canvas, flips that order. Built by Kevin Yang, the tool lets designers sketch a space in the browser the way a web designer blocks out a layout with gray boxes before picking colors. Place a sound source and a listener, open a door or change a wall material, and hear the difference instantly — what Yang calls an "acoustic wireframe."
Deterministic systems handle the geometry and acoustic calculations, and audio renders locally in the browser. GPT-5.6's role is limited to composing and describing scenes within a constrained schema.
"Shooting more rays doesn't make a better product."
Second place, Sentinel, is a security scanner for MCP servers — the servers that give agents access to files, APIs, databases, and shell commands. Malik Bashaar Javaid kept running into MCP servers shared as "quickstart templates" that left raw shell calls, hardcoded credentials, and loose authorization boundaries wide open.
Sentinel combines deterministic static analysis, a tightly constrained GPT-5.6 review that only operates within actual source context, and Docker-isolated probes. The model can support or challenge a finding, but it can't invent an execution probe or cite code that doesn't exist. Results map to the OWASP Agentic Top 10 and can feed directly into GitHub code scanning.
"Constraining the model was harder than writing the prompt." This was his first hackathon, fresh out of a computer science degree.
Three finalists rounded out the category. Emberframe Studio is a spatial canvas for building with AI agents. GenUI converts model-written JSON into verified native SwiftUI. Vibe Signal lets you monitor and direct Codex from an iPhone or Apple Watch.
Education — breaking through after 25 years as a developer
First place, Mechanica, began in a museum. "I don't know how many times I've bumped my head against the glass — I got so absorbed in what was behind it," says Shan Wei. Looking wasn't enough. She wanted to touch it, take it apart.
The project revives four ancient Chinese machines — known today from only a few surviving lines of text — as working reproductions, ranging from an astronomical clock tower to a programmable loom that predates modern computing by centuries. Visitors can spin the machines apart and trace every dimension back to classical texts, surviving artifacts, or clearly labeled scholarly inference. Where scholars disagree, the project shows competing reconstructions side by side instead of picking one answer.
Weiying Zhu, Yukun Li, and Shan Wei all had long software careers but had never touched 3D modeling, physics simulation, or animation before this. Codex carried them across that unfamiliar territory, and GPT-5.6 became an AI docent that cites the museum's sourcing — or declines to answer when there isn't one.
"I've built software for 25 years, and 3D and animation were completely outside my wheelhouse — but that wall just disappeared."
Second place, Dấu, was built by Robert Huynh, who once walked out of a hackathon halfway through, convinced his idea wasn't "technical enough." This time he showed up solo, from Hanoi.
In Vietnamese, the same syllable can mean six different things depending on pitch and voice quality. Dấu lets learners record a word, then plots their own pitch curve next to a verified native-speaker reference, tells them what meaning their pronunciation actually produced, and gives specific physical corrections. Tone detection runs on deterministic signal processing, while GPT-5.6 handles only the coaching. When the evidence is unclear, it asks the learner to try again rather than confidently guessing wrong.
"I went from wondering if I was technical enough to even enter a hackathon, to building something solo that made it to finalist."
Three finalists rounded out the category. Canopy lets learners apply what an AI just taught them directly in a coding sandbox. Encore! turns a kid's wrong test answer into a custom picture-book lesson. ResearchOS traces every research claim back to its source.
Winners at a glance
| Category | First Place | Second Place |
|---|---|---|
| Apps for Your Life | Second Voice — turning fragmented speech into sentences | AirBridge — streaming Windows audio to AirPlay |
| Work & Productivity | veTriage — phone triage for a veterinary clinic | Pulse — tracking cardiac arrest resuscitation |
| Developer Tools | Echo Canvas — acoustic wireframing | Sentinel — MCP server security scanner |
| Education | Mechanica — reconstructing ancient Chinese machines | Dấu — a Vietnamese tone-pronunciation coach |
Editor's take
Line up all eight winners and two patterns jump out.
First, the winners weren't the best coders — they were the people who understood the problem. Two of the four first-place teams are a veterinarian and a cardiologist. The other two crossed into fields outside their own expertise — 3D modeling and spatial audio — using Codex to get there. Hackathons used to be a contest of raw build speed. This one tilted toward domain knowledge instead.
Second — and this matters more for practitioners — most of the winning entries kept AI on a short leash.
Pulse handed resuscitation-protocol tracking to deterministic code and gave the model only the job of interpreting speech. Echo Canvas built its own geometry and real-time audio engine and used the model purely as an authoring and explanation layer. Dấu ran tone detection through signal processing and gave the model nothing but coaching duties. The builder of Sentinel said outright that most of the work wasn't prompting — it was building constraints.
Anyone who's spent real time with coding agents knows exactly where this usually falls apart. The wider the scope you hand the model, the flashier the first demo looks — and the shakier the results get by the second run. In practice, that means most of the real engineering time goes into narrowing what the model is allowed to touch, using schemas and hard rules.
That's precisely what these winning projects have in common. Judgment, computation, and procedure stay in human-written code; the model only gets put in charge of understanding and explaining what people say. And when the evidence is thin, it asks rather than guesses.
The fact that eight very different projects, chosen out of 8,000 entries, converged on the same approach isn't a matter of judges' taste. It looks more like an answer to the question of how this technology should actually be used right now: the model earns its keep as an adapter, not as the engine.
The same logic applies directly to teams here. Organizations that understand their own ground truth — hospitals, warehouses, legal departments, manufacturing floors — are in the best position right now. What they need isn't more engineering headcount. It's putting Codex in the hands of someone with 20 years on the floor and giving them two clear weeks. veTriage was built exactly that way, and it's running in a real clinic today.





Comments