Source description
About the role
Implement and evolve voice agent capabilities, including real‑time speech processing, multimodal orchestration, and agent runtime integration Develop and integrate talking avatar technologies, supporting both zero‑shot and customized experiences Apply and adapt modern generative AI technologies (including diffusion and language modeling) where appropriate, with an emphasis on engineering robustness and system integration Collaborate with cross‑functional teams to integrate voice agents, speech services, and avatar components into platforms, SDKs, and end‑to‑end applications Continuously improve system latency, quality, scalability, and reliability in real‑world deployments Stay current with industry and research advancements in voice AI, agentic systems, and generative technologies, and translate them into practical solutions Contribute to engineering excellence through design reviews, code reviews, and shared best practices Bachelor's Degree in Computer Science or a related technical field AND 2+ years of professional software engineering experience Strong coding skills in one or more languages such as C, C++, C#, Java, JavaScript, or Python Experience working with machine learning-powered systems, platforms, or services (model development experience is a plus but not required) Background in speech processing, natural language processing (NLP), computer vision, or real‑time systems Familiarity with generative AI technologies (e.g., diffusion, autoregressive approaches), model integration, or optimization Experience building or integrating voice agents, AI services, or large‑scale distributed systems Strong problem‑solving skills with a systems‑thinking mindset Effective communication skills and experience working in cross‑functional teams
More at Microsoft