My research focuses on multimodal foundation models, spatial acoustic reasoning, and generative audio architectures.

For complete citation indexes, visit my Google Scholar Profile and Google Research Profile.


Foundation Models & Publications

2026

  • PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
    Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
    International Conference on Machine Learning (ICML), 2026

2025

  • SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
    Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Vivek Kumar
    arXiv Preprint, 2025

2023

  • SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
    Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin P. Murphy, Alexander G. Hauptmann, Lu Jiang
    Neural Information Processing Systems (NeurIPS), 2023

Audio AI: Challenges, Breakthroughs & Applications

PyTorch DevCon

Detailed walkthrough of deep learning architectures applied to audio signals, comparing classical DSP against learned representations.


Selected Patents