One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

ElevenLabs adds FLUX 3, Seedance 2.5 to creative tools

From new image and video generation models to audiobook character casting and music vocal generation, here's a rundown of August's updates

이미지: YouTube 영상 갈무리

Summary

  • ElevenLabs added Black Forest Labs' FLUX 3 and ByteDance's Seedance 2.5 to its ElevenCreative image and video tools
  • New features include Character Casting, which suggests voices by distinguishing characters when an audiobook manuscript is uploaded, and Spotlight, which scores support conversations in real time
  • ElevenMusic now supports vocal generation, and the Summer of Sound contest, running through August 31, raised the free-tier monthly generation limit to 400
ElevenLabs Changelog: Everything We Shipped This Month
ElevenCreative 신규 모델
FLUX 3(Black Forest Labs), Seedance 2.5(ByteDance)가 이미지·영상 도구에 추가됨
Character Casting
오디오북 원고 업로드 시 등장인물 자동 식별, 목소리 제안, 실제 대사로 미리듣기 지원
Spotlight
ElevenAgents의 통화·채팅 실시간 관찰, 주제 분류, 사용자 지정 기준별 점수화
지원 채널 확장
SMS·Telegram·Intercom·Freshdesk가 기존 전화·웹·Zendesk·Slack·WhatsApp에 추가
ElevenMusic 보컬
자신의 목소리 또는 보컬 라이브러리로 보컬이 들어간 곡 생성 가능
Summer of Sound
8월 31일까지 진행, 무료 이용자 월 400회 생성, 상위 10곡 대상 5만 달러 상금
Dubbing v2
대본이 아닌 원본 연기를 기준 삼아 톤·감정을 유지한 더빙 API
Scribe v2 Realtime
개체 인식, 동일 스트림 내 보조 언어 인식, 배경음 필터링, 무보관 로깅 추가

Upload an entire audiobook manuscript, and the AI will automatically identify the characters and assign voices to them. It suggests voice candidates suited to each character, letting you preview them with actual dialogue before finalizing your choice.

That's the new "Character Casting" feature featured in ElevenLabs' August changelog video posted on YouTube. Following the general availability launch of its real-time conversational voice model "v3 Conversational" on August 20, the company has, within a single month, expanded its reach into image and video generation tools, support agents, music generation, and dubbing APIs. A single video running just over five minutes covers as many as eight updates.

Two external models join the image and video tools

The image and video generation features within ElevenCreative, the company's video content creation tool, now include Black Forest Labs' FLUX 3 and ByteDance's Seedance 2.5. Both are image and video generation models built by outside developers, which ElevenLabs has integrated into its own tool as selectable options.

The audiobook production feature gained Character Casting. Upload a manuscript, and it automatically distinguishes characters and suggests a voice for each one, letting you preview actual dialogue before finalizing your choice.

Spotlight scores support conversations in real time

ElevenAgents, the company's customer support agent product, gained a new observation layer called Spotlight. It reads calls and chat conversations in real time, groups them by topic, scores each conversation against criteria written in plain sentences by the user, and flags what needs to be fixed.

Supported channels also expanded. SMS, Telegram, Intercom, and Freshdesk joined the existing lineup of phone, web, Zendesk, Slack, and WhatsApp, and behavior can now be configured separately for each channel.

ElevenMusic adds vocals to songs

ElevenMusic, the music generation tool, now supports vocals. Users can select their own voice or one from the vocal library to create songs with consistent vocals throughout. The References feature lets users upload a track between 10 seconds and 5 minutes long and generate a new song in that style.

The "Summer of Sound" contest, running through August 31, raised the monthly generation limit for free-tier users to 400 and will split a $50,000 prize pool among the ten most-streamed tracks. To enter, users need to upload a song on the Summer of Sound entry page before August 31.

Dubbing and real-time transcription also got upgrades

ElevenAPI, the developer-facing API, added Dubbing v2. While the previous version dubbed based on the script, this version bases dubbing on the original performance itself, so tone and emotion survive translation into other languages rather than being flattened out.

Scribe v2 Realtime, the real-time speech recognition feature, gained entity recognition, the ability to recognize a secondary language within a single stream, background noise filtering, and a zero-retention logging option that doesn't store data.

This month's updates at a glance

Product AreaNew FeatureKey Details
ElevenCreativeFLUX 3 / Seedance 2.5Choose between external image/video models within the tool
ElevenCreativeCharacter CastingAutomatic voice suggestions for each audiobook character
ElevenAgentsSpotlightReal-time observation and scoring of support conversations
ElevenAgentsChannel expansionAdded SMS, Telegram, Intercom, Freshdesk
ElevenMusicVocal generation / ReferencesGenerate songs with vocals, reference-based styling
ElevenAPIDubbing v2Dubbing based on original performance
ElevenAPIScribe v2 RealtimeEntity recognition, multilingual support, zero-retention logging

How to try it

Image, video, and audiobook creation all happen within ElevenCreative. After logging into an ElevenLabs account, users can access the image tool, video tool, and audiobook tool within ElevenCreative.

For audiobooks, the steps are as follows: 1. Upload a manuscript file to the audiobook tool. 2. Character Casting analyzes the manuscript and displays a list of characters. 3. Review the suggested voice candidates for each character. 4. Preview actual dialogue and finalize the voice you like. 5. If images or video are needed, select FLUX 3 or Seedance 2.5 from the same screen and enter a prompt.

The announcement did not specify availability terms by country or pricing tier. However, the Summer of Sound event is open to free-tier users, who can generate up to 400 songs per month.

For example, an audiobook producer could automatically get character-specific voice assignments from a single manuscript, cutting down on recording preparation time. A call center operations team could use Spotlight to score support quality against defined criteria without having to listen to every call manually. A musician could upload a favorite track to References and layer their own voice onto a new song with a similar mood.

Editor's take

What makes this changelog interesting is that ElevenLabs makes no effort to hide the fact that it doesn't build every model itself. FLUX 3 and Seedance 2.5 were built by Black Forest Labs and ByteDance respectively, and ElevenLabs has simply embedded them into its own creative tools, letting users choose between them on a single screen. Expanding from its original domain of voice into images and video this way looks more like assembly than in-house development. Given that this batch of assembly-style updates arrived just days after the GA launch of the real-time conversational model v3 Conversational on August 20, the company appears to be positioning itself not around voice alone, but around bringing text, voice, image, and video together under a single account.

For teams already using voice AI in production, Spotlight is probably the update worth watching most closely. Support quality management has traditionally relied on someone re-listening to calls and filling out a checklist; if that shifts to a system where writing criteria in plain sentences yields a score the moment a conversation ends, QA staff could be freed up for other work. But adding channels doesn't automatically improve scoring accuracy. Whether the same criteria hold up on short-message channels like SMS or Telegram is something that will only become clear through actual use.

For audiobook producers and independent creators, Character Casting alone lowers the barrier to turning a manuscript into voiced content without a recording studio. Next month's changelog will likely reveal how much the newly added FLUX 3/Seedance 2.5 combination actually gets used, and what default metrics Spotlight settles on for scoring support conversations.

Comments