METAL

DeepMind rebuilt a day that was never filmed

No photograph or film survives of the day Burt and Ethelle Shatz first met. More than 70 years into their marriage, with his memory fading, Google DeepMind mapped the couple's present-day mannerisms onto their younger faces to rebuild that day.

DeepMind rebuilt a day that was never filmed

Summary

  • Google DeepMind and Primordial Soup have released Love, Rendered, a documentary short that reconstructs the day Burt and Ethelle Shatz first met, without a single photograph of it.
  • The method has two stages. Generative models restored black-and-white photographs of the couple's youth to anchor the likeness, and performance capture models mapped their present-day micro-mannerisms onto those younger faces.
  • Ethelle joined as a co-creator, correcting the curve of a staircase and the shape of a shoe heel. Michael Chang, a Google DeepMind engineer, wrote that the team generated a memory the couple told them felt authentic.
Recreating a 70-year love story frame by frame

Rebuilding a day that was never filmed was not a problem of making a convincing picture. Burt and Ethelle Shatz have been married for more than 70 years, and the memory they hold most dearly is the day they met at a student co-op in Cleveland. That day was never photographed or filmed, so it existed only in their heads. As Burt's cognitive decline set in, even that was beginning to thin out.

Google DeepMind and Primordial Soup, Darren Aronofsky's creative venture, rebuilt the day. The result is a documentary short called Love, Rendered. Academy Award nominee Liz Garbus directed it, and Dan Cogan and Aronofsky produced alongside her. The company published an account of the process on September 9, and two days later the DeepMind account posted a 75-second production video.

Michael Chang, a Google DeepMind engineer, served as technical lead. He wrote that he knew going in the work would be as emotionally demanding as it was technically hard, because losing memory is personal to him. On his last visit to a grandfather who had suffered a stroke and memory loss, Chang was already in his thirties, and his grandfather, still certain he was in college, congratulated him on graduating.

That experience became the starting point. When the project began, Chang asked his father for old family photographs, tested the restoration on pictures from around the time his parents met, and used video models to set them in motion. Watching his parents move as twenty-somethings convinced him these tools could hold on to what was slipping away.

The method splits into two stages. First, generative models restored black-and-white photographs of the couple in their youth. Those restored pictures became the anchor that kept every later scene from drifting away from the real people. Then performance capture models carried the small habits of the present-day Burt and Ethelle onto their younger faces: the particular angle at which Burt tilts his head, the brief hesitation in the middle of a sentence, the way the skin folds at the corners of his eyes.

Chang wrote on the blog, "Combining these two mediums allowed us to intertwine Burt and Ethelle's past and their present." He added, "In doing so, we generated a 'memory' that they told us felt authentic." Putting that word in quotation marks himself was his way of marking that what he built is not a record but a generated thing.

In the 75-second production video METAL reviewed, the team explains more precisely where the difficulty sat. The task, they say, was not merely to generate pixels but to recreate a moment that mattered to this family and was never captured in photographs or video. They knew from the outset that simply making a face resemble a person would not produce an emotional or real connection. It could not be just any smile; it had to be details that differ from person to person, such as the way someone tilts their head while listening.

The detail most worth watching in this project is where Ethelle sat. According to the company, she sat beside the team as an active co-creator, correcting the curve of a staircase or the shape of a shoe heel. In the production video the team says the creative team and the film's subjects were the ones who pointed out whether the technology was capturing what actually mattered. The person being depicted stayed in the room throughout, holding the power to change the result.

That arrangement is what separates this work from ordinary face synthesis. In rebuilding a memory, what really divides the cases is not the precision of the technology but who remains the author of that memory. Here the owner of the memory supplied the material, reviewed the scenes that came back, and judged what matched her own. Authorship of the rebuilt day stayed with the person who lived it rather than with the model.

Garbus and Aronofsky each came to the subject for their own reasons. While making the documentary Coma, Garbus watched fMRI scans light up when patients in minimally conscious states heard familiar voices or were shown pictures of people they loved. Aronofsky had seen footage of a former ballerina with Alzheimer's who, on hearing Swan Lake, danced the choreography from her wheelchair. Those two experiences led the team to reminiscence therapy, the clinical practice of using sensory cues such as songs and old photographs to wake memories.

The question Love, Rendered raises sits in the next square over. Reminiscence therapy works when a cue survives to do the waking, and for Burt and Ethelle's day there was no cue at all. So instead of finding one, the team made one. The company says Aronofsky remarked during production that a tool, like a paintbrush or a hammer, does nothing until a human hand guides it.

METAL has reported on Higgsfield releasing a tool that keeps the performance and the camera move while swapping out the background, and lifting human motion to carry it elsewhere is already a product. What makes this work different is one thing: the person who supplied that motion approved the result from off camera. Used without the subject present, the same technology becomes the case on the opposite side.

Building models from material the owner agreed to is taking hold elsewhere too. METAL has covered Suno retraining its models on licensed catalogs from Warner and BMG, and generative tools are moving toward disclosing where their material comes from. With likeness, that material is a person's face and habits, so the standard has to be stricter.

Google has opened the same restoration capability in the Gemini app. Upload an old photograph, ask for restoration and colorization, and tell it to preserve the appearance, expression and pose of the people. The performance capture used in the film is not part of that release; what is public is the photo restoration.

What this work leaves behind is a constructed scene, not a record. DeepMind did not hide that, putting the word memory in quotation marks, and it handed the verdict on authenticity not to the technology but to the two people who lived the day. This short film has shown, ahead of the rest, the order that generative technology has to follow when it handles a human face.

Comments