Apply at Tencent UK More roles at Tencent UK
1. Research and develop advanced speech synthesis and generation algorithms (e.g., TTS, voice conversion, sound/music generation) based on LLMs, multimodal/omnimodal LLMs.
2. Develop and optimize speech and audio synthesis systems for online applications, improving effectiveness, efficiency, and scalability.
3. Explore and advance full-duplex/streaming multimodal LLM capabilities in speech understanding, generation, and real-time spoken interaction.
4. Collaborate cross-functionally with research and engineering teams from prototyping to production.
1. Ph.D. in Computer Science, Electrical Engineering, Signal Processing, or a closely related field.
2. Strong foundation in speech/audio processing and modern generative models (e.g., diffusion, flow matching, autoregressive, codec-based approaches).
3. Hands-on experience extending LLMs to speech/audio modalities (e.g., speech tokenizers, multimodal adapters, speech-text joint training). Experience with full-duplex or streaming spoken dialogue systems and real-time interaction modeling is a plus.
4. Proficient in Python and deep learning frameworks (e.g., PyTorch); experience with distributed training is a plus.
5. Track record of publications at top-tier venues (e.g., ICML, NeurIPS, ICLR, ACL, ICASSP, Interspeech).
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.