real-time agents · full-duplex modeling · multimodal agents
Hi 👋, I’m Weiting (Steven) Tan, a Research Scientist at ByteDance Seed working on real-time agents. I completed my PhD in Computer Science at Johns Hopkins University’s Center for Language and Speech Processing, advised by Prof. Philipp Koehn. I also completed my undergraduate and master’s degrees in Computer Science at JHU. Go Hop!
My research focuses on real-time agents, full-duplex modeling, and multimodal agents. I am interested in systems that can listen, speak, perceive, reason, and act continuously—handling overlapping turns, low-latency interaction, and multiple modalities within one coherent agent.
Recently, I have been working on real-time multimodal interaction, tool-using agents, and unified models for multimodal understanding and generation. I wrote about my research journey—and where I think voice agents need to go—here.
news
| Apr 19, 2026 | I defended my PhD at Johns Hopkins University! |
|---|---|
| Apr 06, 2026 | Starting a full-time Research Scientist role at ByteDance Seed, working on real-time agents. |
| May 01, 2025 | Starting a new internship at Bytedance Seed! Will be working on iterative (multi-turn, multi-modal) tool-use agents. |
| Aug 02, 2024 | I will work on Multi-modal LLMs as a part-time student researcher at Meta in Fall 2024 & Spring 2025. |
| Feb 20, 2024 | I will be interning at Meta AI (FAIR) in summer 2024, working on Speech Large Language Model. Looking forward to the new project! |
| Apr 07, 2023 | I will be staying at Johns Hopkins University for my PhD, working with Prof. Philipp Koehn! |
latest posts
| Aug 15, 2026 | Beyond Full Duplex: The Voice Agent as an Event Loop |
|---|---|
| Aug 15, 2026 | 超越全双工:把语音智能体理解为事件循环 |
| Mar 29, 2026 | 因果: Reflections on a PhD Journey |
selected publications
- arXivFlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation2026
- arXivProcess-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents2025
- EMNLPSeeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation2025
- IWSLT
- NeurIPSDiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation2024
- IWSLT