Staff Research Engineer - Multimodal Generative Modelling
- Hiring Organisation
- Jobleads-UK
- Location
- Greater London, England, United Kingdom
short and long time horizons. Propose novel multi-modal system architectures (especially text and voice). Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis. Design solutions that reinforce emotional expressiveness and natural interaction. Implement and bring designs to life, from pretraining through … training. Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness. Define new evaluation metrics for conversational systems, including latency‐aware and interaction‐based measurements. Track the latest research in audio‐visual diffusion, autoregressive models, neural codecs, and multimodal LLMs. Curate new datasets ...