Weiting Tan

Center for Language and Speech Processing, Johns Hopkins University

prof_pic.jpg

real-time agents · full-duplex modeling · multimodal agents

Hi 👋, I’m Weiting (Steven) Tan, a Research Scientist at ByteDance Seed working on real-time agents. I completed my PhD in Computer Science at Johns Hopkins University’s Center for Language and Speech Processing, advised by Prof. Philipp Koehn. I also completed my undergraduate and master’s degrees in Computer Science at JHU. Go Hop!

My research focuses on real-time agents, full-duplex modeling, and multimodal agents. I am interested in systems that can listen, speak, perceive, reason, and act continuously—handling overlapping turns, low-latency interaction, and multiple modalities within one coherent agent.

Recently, I have been working on real-time multimodal interaction, tool-using agents, and unified models for multimodal understanding and generation. I wrote about my research journey—and where I think voice agents need to go—here.

real-time agents full-duplex modeling multimodal agents

news

Apr 19, 2026 I defended my PhD at Johns Hopkins University!
Apr 06, 2026 Starting a full-time Research Scientist role at ByteDance Seed, working on real-time agents.
May 01, 2025 Starting a new internship at Bytedance Seed! Will be working on iterative (multi-turn, multi-modal) tool-use agents.
Aug 02, 2024 I will work on Multi-modal LLMs as a part-time student researcher at Meta in Fall 2024 & Spring 2025.
Feb 20, 2024 I will be interning at Meta AI (FAIR) in summer 2024, working on Speech Large Language Model. Looking forward to the new project!
Apr 07, 2023 I will be staying at Johns Hopkins University for my PhD, working with Prof. Philipp Koehn!

latest posts

selected publications

  1. arXiv
    FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
    Weiting Tan, Andy T. Liu, Ming Tu, and 3 more authors
    2026
  2. arXiv
    Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
    Weiting Tan, Xinghua Qu, Ming Tu, and 4 more authors
    2025
  3. EMNLP
    Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
    Weiting Tan, Jiachen Lian, Hirofumi Inaguma, and 3 more authors
    2025
  4. IWSLT
    SSR: Alignment-Aware Modality Connector for Speech Language Models
    Weiting Tan, Hirofumi Inaguma, Ning Dong, and 2 more authors
    2024
  5. NeurIPS
    DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
    Weiting Tan, Jingyu Zhang, Lingfeng Shen, and 2 more authors
    2024
  6. IWSLT
    Streaming Sequence Transduction through Dynamic Compression
    Weiting Tan, Yunmo Chen, Tongfei Chen, and 5 more authors
    2024