I work on robust “speech-in-the-wild” processing: overlapping speakers, noise, reverberation, and distant microphones. My research spans the full speech pipeline — separation, enhancement, speaker diarization, and automatic speech recognition — with a focus on end-to-end integration and foundation models that natively understand multi-talker conversational speech. I also work on text-to-speech and speech generation for synthetic data creation, enabling speech-in-the-wild systems in domains where real data is scarce, such as doctor–patient conversations.
I am currently a postdoctoral researcher in Prof. Shinji Watanabe's WAVLab at Carnegie Mellon University. I received my PhD in Information Engineering (2023) and MSc in Electronic Engineering (summa cum laude, 2019) from Università Politecnica delle Marche, advised by Prof. Stefano Squartini. Along the way I have worked with Amazon Alexa (applied scientist intern), consulted for startups building AI-powered hearing aids (Fortell, formerly Chromatic, which has since raised $163M at a $740M valuation) and silent speech interfaces (AlterEgo), and took part in three JHU JSALT workshops — as a participant in 2019 (during my master's) and 2026, and leading, with Lukáš Burget, the JSALT 2025 EMMA team on end-to-end multi-channel multi-talker ASR, which has produced six papers to date.
I believe strongly in open and reproducible science: I co-authored and contribute to several widely used open-source speech toolkits (SpeechBrain, ESPnet, Asteroid, TorchAudio — collectively exceeding 23,000 GitHub stars), and I organize community challenges — CHiME-7/8 DASR, URGENT, DCASE Task 4 — that push the field toward robustness and generalization. Core architectures I co-developed, such as SepFormer and TF-GridNet, are widely adopted by industrial research labs — including Google, Microsoft, MERL, and Samsung — for speech separation and enhancement. I served as secretary of the ISCA Special Interest Group on Robust Speech Processing (2023–2025) and serve as Area Chair / meta-reviewer for ICASSP 2026, Interspeech 2026, and IJCNN 2025–26, with 60+ publications and 4,000+ citations in these areas.
When I am not doing research, I like to hike, ski (picture above), swim or — very recently — play tennis and disc golf (I am bad at both; I used to be decent at soccer though). I used to play the electric guitar; now I only play the keyboard (only the laptop one, sadly).
News
- Aug 2026Took part in the JSALT 2026 workshop as a member of the Spoken Conversational AI Simulator team, on simulating full-duplex conversations for evaluating AI systems.
- Sep 2026Together with colleagues from NTT and Brno University of Technology, we will give a tutorial on Conversational Speech Recognition and Analysis: Progress, Challenges, and Emerging Opportunities with Speech Language Models at Interspeech 2026 in Sydney (28 September – 1 October 2026).
- Sep 2026Eight papers accepted at Interspeech 2026.
- Fall 2026Teaching CMU's 11-751/18-781 Speech Recognition and Understanding course.
- May 2026Invited talk at the JHU Human Language Technology Center of Excellence (HLTCOE).
- Feb 2026Invited talks at the CMU LTI Colloquium and the MILA Conversational AI reading group.
- Jan 2026MAPSS accepted at ICLR 2026; SE-DiCoW and the URGENT challenge accepted at ICASSP 2026. Invited talk at MBZUAI.
- Jan 2026Five papers from the JSALT 2025 EMMA team submitted to ICASSP 2026 — new state-of-the-art results on several multi-talker ASR benchmarks and a new multi-talker ASR framework based on shuffle automata.
- Dec 2025ASRU 2025 hackathon prize for best technical quality and correctness. ARECHO presented at NeurIPS 2025.
- Oct 2025Invited talks at Microsoft and the CMU Speech Lunch.
- Aug 2025Wrapped up the JSALT 2025 workshop in Brno, where I led with Lukáš Burget the 17-person EMMA team on end-to-end multi-channel multi-talker ASR — an effort that has since produced six papers.
- Jun 2025OpenBEATs (WASPAA 2025), ESPnet-SpeechLM (NAACL 2025), CS-FLEURS (Interspeech 2025) and the Interspeech 2025 URGENT challenge papers are out.
- 2024Awarded a Meta Audiobox Research Grant. Guest editor for the Computer Speech & Language special issue on multi-speaker, multi-microphone distant speech recognition.
Open-Source Software
- SpeechBrain — core developer and co-author of one of the most used PyTorch speech toolkits (10k+ stars).
- ESPnet — top-10 all-time contributor; co-author of ESPnet-SE++, ESPnet3 and the ESPnet-SpeechLM toolkit.
- Asteroid — co-author of the PyTorch audio source separation toolkit, natively integrated on the Hugging Face Hub (340+ models) and at one point featured on the Hugging Face landing page alongside organizations like fairseq and ESPnet.
- TorchAudio — co-author of TorchAudio 2.1.
- Contributor to Lhotse, DCASE-REPO, NVIDIA NeMo, Kaldi, and Pyroomacoustic.
Selected honors
- ASRU 2025 hackathon prize for best technical quality and correctness
- 1st place, 2nd Clarity Enhancement Challenge (2022) — demo
- 1st place, L3DAS22 Speech Enhancement Challenge (2022)
- 1st place (Developer Tools track), Global PyTorch Summer Hackathon 2020 — DeMask, built with Asteroid
- Judges' award for most innovative system, DCASE 2020 Task 4 Challenge
- Best student paper (co-author), IEEE SLT 2022
- Meta Audiobox Research Grant (2024)
- Outstanding Reviewer, IEEE ICASSP 2023
Press
- Quoted as an expert in MIT Technology Review: “A new AI translation system for headphones clones multiple voices simultaneously” (May 2025)
- Quoted as an expert in MIT Technology Review: “Noise-canceling headphones use AI to let a single voice through” (May 2024)
Publications
Selected from 60+ publications — see Google Scholar for the full list.
Most cited
Recent
Contact
samuele.cornell [at] ieee.org
Language Technologies Institute, Carnegie Mellon University
5000 Forbes Avenue, Pittsburgh, PA
Always happy to chat about robust speech processing, open-source, challenges, or potential collaborations.