Accent TTS Demo

Joycent: Zero-Shot Accent TTS with LLM-Assisted Accent Embeddings and Conditional Layer Normalization

Abstract

Text-to-speech (TTS) systems have advanced in producing natural speech across speakers and styles, yet zero-shot accent control remains challenging due to the entanglement of speaker and accent attributes. We propose Joycent, a novel accent TTS model that integrates both audio-based and Large Language Model (LLM)-assisted accent embeddings to achieve disentangled and natural accent synthesis. Audio-based embeddings are derived from speech using the pre-trained accent identification model WhisAID, while LLM-assisted embeddings are generated from natural language descriptions of the target accent prompt in text. Speaker embeddings are separately extracted and are conditioned together with the accent embeddings via conditional layer normalization (CLN) modules within the text encoder. Experiments demonstrate that Joycent outperforms baselines in naturalness, accentedness, and speaker similarity, showing effectiveness in zero-shot accent generation without compromising speaker identity. The audio samples are available at https://oshindow.github.io/joycent/.

Model Architecture

Model Diagram

Figure 1. Overall model architecture.

Seen Speaker Samples (SSB0693)

Target Accent Prompt Accent Prompt Speaker MacST Joycent w/o spk CLN w/o LLM emb
"但是争取好成绩的前提是身体好"
Singaporean
Southern -
"我想只有这样才能平息这个风波"
Singaporean
Southern -
"这主要是因为对于很多车主来说"
Singaporean
Southern -
"他带领中国女排闯进决赛"
Singaporean
Southern -
"哦那不错诶你的中文还可以读那些书"
Singaporean
Southern -
"没有我最近呃最近是在看那个密室大逃脱但是已经完了到最后一集我还没看"
Singaporean
Southern -

Unseen Speaker Samples (G0003/SSB1340)

Target Accent Prompt Accent Prompt Speaker MacST Joycent w/o spk CLN w/o LLM emb
"但是争取好成绩的前提是身体好"
Singaporean
Southern -
"我想只有这样才能平息这个风波"
Singaporean
Southern -
"这主要是因为对于很多车主来说"
Singaporean
Southern -
"他带领中国女排闯进决赛"
Singaporean
Northern -
"哦那不错诶你的中文还可以读那些书"
Singaporean
Northern -
"没有我最近呃最近是在看那个密室大逃脱但是已经完了到最后一集我还没看"
Singaporean
Northern -