
Summary
- Higgsfield released "The Curlyhill Boys," a 110-minute AI feature film starring real actors, and detailed the production process in a thread
- UFC fighter Israel Adesanya, Quinton "Rampage" Jackson, and streamer N3on appeared in the film after signing likeness and voice rights agreements
- The contracts covered pay, usage scope, and script approval rights, and required source material to be deleted within 30 days of filming and never used for training
Show more
- The film ran on ByteDance Seed's Seedance model through the Higgsfield platform, and all prompts and assets were made public
A 110-Minute AI Film Starring Real Actors
Higgsfield posted the feature film "The Curlyhill Boys" to its X account on August 10 local time. The production reportedly cost around $2 million, and the video generation ran on ByteDance Seed's Seedance model through the Higgsfield platform. Higgsfield builds its own models, Soul and DoP, while also drawing on models from OpenAI, Google, and ByteDance — and for this project, it chose Seedance to render real actors' faces and voices onto the screen.
The story itself is a crime comedy. Three unknown rappers in East London set out to shoot a music video, end up stealing a boat loaded with cash, and get pulled into a chase between rival criminal gangs. Higgsfield lists the runtime at 110 minutes, though the video posted to X shows a playback length of 1 hour, 54 minutes, and 51 seconds. According to the US content outlet Bottle Rocket, the film premiered at Glasshouse in New York on August 5 before going up online.
The cast includes UFC fighter Israel Adesanya (Stylebender), former UFC champion Quinton "Rampage" Jackson, streamer N3on, and streetball creator MKIATPIS. The thread spells out exactly who plays whom: Oli is Adesanya, Tobin is Rampage Jackson, and Horace is N3on. Higgsfield said it digitized each person's real face through a contracted photo shoot, and that likeness and voice rights agreements were finalized before any generation began.


Three Things Written Into the Likeness Contracts
According to a rundown from the AI industry newsletter AlphaSignal, the contracts rested on three pillars. Pay, usage scope, and script approval rights were all put in writing. The original photos and recordings used to build each face and voice had to be permanently deleted within 30 days of the end of filming. And Higgsfield was barred from retaining that data, using it for training, or relicensing it to third parties. It's a structure that mirrors the advance-notice and compensation provisions SAG-AFTRA won through its strike.
The script followed a similar model. Bottle Rocket reports that screenwriter Timothy Flannagan (outlets spell the name differently, as Flanagan or Planagan) was paid at rates matching Writers Guild of America standards despite not being a guild member. He retained traditional film adaptation rights while licensing only the generative-AI adaptation rights to Higgsfield. The underlying script, five years in development, is a finalist for the 2026 London Independent Screenplay competition. Higgsfield had previously drawn criticism for making "Hell Grind" without negotiating any rights with its cast at all.


Locking Faces Down With Three-Image Character Sheets
According to the production team's disclosed method, each character's sheet is built from three images: a front full-body shot, a back full-body shot, and a 3/4-angle close-up portrait. Faces are erased entirely from the full-body shots. That's because in wide shots the face is small and blurry, and the model tends to copy that blur directly rather than reconstruct a clean face. The rule was to pull faces only from the close-up portrait — and even that portrait had to be made in two versions, one smiling and one neutral. Without that split, the team said, the model tends to invent teeth on its own, making the mouth look like someone else's. Once an asset was finalized, the team stress-tested it by generating it ten more times in different poses and lighting to confirm it stayed recognizable across all ten.
Every change in a lead character's physical state also got its own separate asset. Cal in a clean jacket with an orange backpack, Cal soaked after falling in the Thames, and Cal in Act Three with a cut eyebrow and a blood-stained white T-shirt were kept as three distinct assets rather than bundled into one with notes. Mixing all three into a single prompt, the team found, caused the blood to fade or the wet clothes to dry out on their own. Higgsfield summed up the philosophy: "Splitting is cheaper than arguing."

What Got Locked Into the Voice and Era Prompts
Voice wasn't treated as an asset but as a precisely worded condition sentence specifying accent, pace, and delivery. That exact sentence, unchanged down to the wording, got pasted into the audio input field every single time the character spoke. Even a small rewording, the team found, widens the range the model samples from and makes the voice drift.
The published canvas shows the voice condition for lead character Cal: "A 25-year-old English man. Smooth, relaxed mid-range baritone; casual, unscripted, and conversational delivery; South London street accent — dropped h's, glottal t, -ing to -in'; nostalgic and warmly reflective, sharing a personal memory with a natural, steady pace." Age, register, delivery, accent, and emotional tone are all packed into a single sentence, with the accent spelled out phonetically.
The film's time period was also locked down as a hard rule: nothing manufactured after 2011 could appear on screen. The team explained that models have a built-in pull toward rendering scenes in a present-day style, so any leeway tends to produce stray details like an extra in a crowd holding an iPhone.

Filming Rules Set Entirely Through Text
For locations, instead of generating a straight-on photo, the team first generated a slow camera pan through an empty space, then used that footage as the reference for the model to fill in the rest of the room. One rule held across the entire film: text-to-video only, with no image ever used as a starting frame. Every scene was built purely from reference assets and text.
A single hallway scene stacked multiple rules on top of each other. The line of dialogue was locked to exactly three words: "Pull it, Oli." Height differences between characters were spelled out explicitly, as in "Horace's eye-line at the level of Cal's mouth." Off-screen events were written as a list of things that must not happen, and the two characters' reactions were instructed to never occur at the same moment.
The same prompt even specified lens and color. "ONE continuous ~12–14s 200mm (FOV ≈12°) tight CLOSE-UP two-shot" locked down shot length, focal length, and field of view in a single line, with reaction beats tagged to exact timecodes at 1.5, 4.0, 6.0, 9.0, and 11.0 seconds. The paragraph ends with three more instructions: "Color 60:30:10. Nothing modern beyond 2011. British spelling." — color ratio, an era cutoff, and even screen-text spelling conventions, all packed into one prompt sheet.
For the singing scenes, the team never asked the model to sing. They recorded the track first, cut it into 12-second segments, fed those in as audio files, and labeled each one "THE TRACK THEY ARE PERFORMING." An earlier label, "audio guide, for sync only," was technically accurate but simply didn't produce usable output. For scenes with sustained physical contact — like a fight inside a car — the model couldn't handle it at all, so the team brought in stunt performers, filmed them in a real car on an iPhone, and fed that footage in as a motion reference.

137 Iterations and Five Guiding Principles
Production generated scenes in batches, changing one line at a time, and logged 137 separate generation attempts. Each log recorded the version, what changed, and the outcome, allowing the team to trace exactly which version had introduced a problem. The team also set a rule that if 10 to 15 attempts still weren't working, the fix wasn't to tweak the wording further but to break the scene down or simplify it.
| Area | Core Rule | Reason |
|---|---|---|
| Faces | Extract faces only from close-up portraits, one smiling and one neutral | Prevents distortion in wide shots |
| Voice | Paste the condition sentence verbatim every time | Rewording causes the voice to drift |
| Era | Ban anything manufactured after 2011 | Models default toward present-day settings |
| Filming | Text-to-video only, no starting frame | Preserves continuity between cuts |
| Singing | Pre-record and insert as 12-second audio blocks | Models can't sing on their own |
| Fighting | Film with stunt performers, use as motion reference | Sustained physical contact can't be generated directly |
The team also released five principles underlying the entire film: lock every asset before starting generation; re-describe everything every time, since the model has no memory; change only one thing at a time; give the model one corner of a room rather than the whole room; and if a scene isn't working, simplify the scene itself rather than the wording.

Four Weeks, 28 People, $1 Million in Compute
According to AlphaSignal's tally, production took four weeks and involved 28 people, nine of whom were directors. Forty percent of the team had no prior experience making AI video before joining the project. The film used roughly 1,000 generated assets and 100 AI-generated locations. Of the $2 million budget, about half — $1 million — went to generation costs, meaning compute.
The toolset wasn't limited to one platform. Ukrainian outlet UA.NEWS reported that the team used Anthropic's Claude for prompt writing, Seedance 2.5 for video and audio generation, and ByteDance Seedream and Google Nano Banana for image editing. That's why Higgsfield's thread opens by instructing readers to first grab the CINEDANCE skill and open the canvas. Character sheets, locations, and prop plates all live on a single canvas, governed by one team rule: "If it's on the canvas but not on the shot list, the canvas wins."

Flaws That Remain
Opinions on the finished film are mixed. UA.NEWS pointed out that on-screen text still looks garbled and character interactions feel weak — flaws typical of generated video that haven't been fully solved. A video editor cited by AlphaSignal noted that frame rates dropped below 30fps in some sections and that close-up faces looked "a little too perfect." That same editor added, though, that for general audiences, "it worked as a first attempt."
Higgsfield said it has published every prompt and asset used in the production on its project page, making the character sheet structure, the phonetic accent notation, and the audio-insertion method all available for anyone to reuse directly. The company also launched a free curriculum aimed at helping traditional production crews transition into AI filmmaking.

Editor's Take
When we watched Higgsfield's short sci-fi comedy "Adiliada," made with Cinema Studio 4, back on August 14, it was just that — a short. "The Curlyhill Boys" is a 110-minute feature, and it licenses the faces and voices of real UFC fighters and streamers through actual contracts. In the span of days, the same company jumped from short to feature, and from fictional characters to contracted likeness rights for real people. What's notable is that this trajectory matches exactly what Runway showed at its AIF 2026 Seoul screening. When we covered that event on August 7, the deciding factor wasn't the technology for maintaining character consistency — it was editorial judgment about what to keep and what to cut. Higgsfield's 137 logged iterations and five principles here read almost like a written record of that same editorial instinct.
From a practical standpoint, the most valuable parts of this thread are the rule to split character sheets into three separate images and pull faces only from close-ups, and the rule to copy-paste voice condition sentences verbatim without changing a word. Anyone who's worked with AI video knows the problem: a character's face shifting subtly from cut to cut usually isn't a flaw in the character sheet itself, but a design mistake in how the face gets extracted from it. The team's method of generating an asset ten more times to confirm it stays recognizable is a verification technique any individual creator could adopt today. Just as instructive is the decision not to force text-only generation in areas the model clearly can't handle yet, like singing or fighting, and instead bring in live-action footage as a motion reference. Right now, the state of the art in AI filmmaking comes down to knowing exactly where the model's limits are.
But separate from all this technical detail, the choice of subject matter deserves attention. The fact that real UFC fighters and streamers licensed their likenesses into a fictional crime narrative is itself a signal that likeness licensing is becoming one axis of film casting. That's exactly why it matters that the 30-day deletion clause, the ban on training use, and script approval rights were all written into the contract. The truly new thing here may not be the prompts at all — it may be this contract template. In the coming weeks, it seems likely other AI video platforms will start replicating the same combination: real celebrity likeness deals paired with open-sourced prompts. If so, it won't be the content itself but this contractual structure and prompt workflow that becomes the industry's next standard.

Sources
- X — 미디어·생성AI — We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000. →
- X — 미디어·생성AI — We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000. →
- X — 미디어·생성AI — We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000. →
- AlphaSignal — Higgsfield Made a $2M AI Film With Licensed Celebrity Likenesses →
- Bottle Rocket Content — Higgsfield's $2M AI Movie Is Free and Open-Sourced →
- UA.NEWS — Higgsfield has released the AI-generated film Cully Hill Boys and prompts for it →





Comments