
- 프로젝트
- audio.cpp
- 버전
- Release 0.5
- 추가 기능
- DramaBox 표현형 TTS · Confucius4 교차언어 음성 전이
- 모델
- 7종 추가 지원
- 백엔드
- ROCm/HIP 지원 포함
audio.cpp 0.5 is out. Expressive TTS (DramaBox) and cross-lingual voice transfer (Confucius4) have been added, along with 7 newly supported models and a ROCm/HIP backend.
This means the local voice stack has gotten one notch thicker. Organizations that need to handle voice data that can't be sent to a cloud API now have more options.
What these two features change
Expressive TTS goes beyond flatly reading text out loud — it carries emotion and inflection. This opens up use cases that plain-sounding synthetic voices couldn't handle before: narration, audiobooks, character voices.
Cross-lingual voice transfer carries the vocal characteristics of a recording made in one language onto speech in another language. This means producing multiple language versions with the same speaker is now possible locally. Its biggest practical use is maintaining speaker consistency across multilingual dubbing.
| Addition | Practical use |
|---|---|
| DramaBox expressive TTS | Narration · audiobooks · character voices |
| Confucius4 cross-lingual voice transfer | Multilingual dubbing, speaker consistency |
| 7 additional supported models | More options across quality, speed, and size |
| ROCm/HIP backend | Local operation on AMD GPUs |

ROCm support is the quiet key point
Local voice stacks have effectively been tied to a specific GPU ecosystem until now. With a ROCm/HIP path open, the same pipeline can also run in environments using AMD cards.
More hardware options also means lower procurement costs. Voice synthesis has smaller memory requirements than large language models, so a high-end consumer card is often enough. Once card brand no longer matters, you can just buy whatever's in stock.
When running locally makes more sense
| Condition | Local | Cloud API |
|---|---|---|
| Original audio can't leave the organization | ● | ✕ |
| Large-scale batch processing | ● | Cost spikes |
| Real-time conversational response needed | Depends on hardware | Network latency |
| Low-volume, intermittent use | Over-investment | ● |
| Always wanting the latest quality | Update burden | ● |
If there's voice data that can't leave the building — internal meeting recordings, consultation transcripts, unreleased content — local is effectively the only answer.
What to check before using it
- Voice rights. Cross-lingual transfer carries over the vocal characteristics of a real speaker. Using someone's voice as a source without their consent creates legal exposure. Even for internal use, it's better to keep consent documentation on file. For voice actor or announcer recordings, separately check whether the contract includes a clause covering synthetic use.
- Disclosure. Labeling requirements for synthetic voice in content differ by platform. Standards are stricter for advertising and political content.
- Licensing. The 7 newly added models may each carry different terms. Whether commercial use is allowed needs to be checked individually on each model card.
- Quality review process. Expressive TTS has a failure mode where inflection gets overdone. Skipping a human listening pass on the final output leads to mistakes.
Summary
The trend of local voice stacks getting thicker continues on its own. What matters right now is less about the technology and more about process. Sort out the rights situation for voice sources and set an internal standard for labeling synthetic output, and it'll carry over even as the tools change.
Source: r/LocalLLaMA release post. See the project repository for detailed specs.