See Google Scholar for a complete listing of papers and patents.


Papers

  • PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs – Dementyev, Zulfikar, Hersek, Getreuer, A. Kumar, V. Kumar – ICML, 2026 [arXiv] [pdf]

  • SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization – Dementyev, Zulfikar, Hersek, V. Kumar – arXiv preprint, 2025 [arXiv] [pdf]

  • SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs – Yu, Cheng, Wang, V. Kumar, Macherey, Huang, Ross, Essa, Bisk, Yang, Murphy, Hauptmann, Jiang – NeurIPS, 2023 [paper] [pdf]

  • Voice conversion with conditional SampleRNN – V. Kumar et al. – Interspeech, 2018 [arXiv] [pdf]

  • Transform-domain decorrelation in Dolby Digital Plus – V. Kumar et al. – ICASSP, 2014 [IEEE Xplore] [pdf]

  • Pseudo-Reliable Code Development – V. Kumar – Embedded Design, 2000 [article] [pdf]


Talks & Presentations

I speak about Audio AI, deep learning for signal processing, and multimodal understanding and generation. My talks cover practical applications of AI in audio, the intersection of machine learning and traditional signal processing, and how foundation models are reshaping how we work with sound, speech, and music.

Audio AI: Challenges, Breakthroughs & Applications

PyTorch DevCon, 2019

Walkthrough of deep learning architectures applied to audio signals, comparing classical DSP against learned representations.

Artificial Intelligence in Audio: Applications, Advancements And Trends

Dolby Soho, New York, 2019

Keynote on emerging trends in audio generation, spatial acoustics, and machine intelligence applications.

Learning Deep Learning

Deep Learning Meetup, 2017

Foundational principles and hands-on guidance for engineering teams transitioning into deep learning architectures.


Patents