I work on robust “speech-in-the-wild” processing: overlapping speakers, noise, reverberation, and distant microphones. My research spans the full speech pipeline — separation, enhancement, speaker diarization, and automatic speech recognition — with a focus on end-to-end integration and foundation models that natively understand multi-talker conversational speech. I also work on text-to-speech and speech generation for synthetic data creation, enabling speech-in-the-wild systems in domains where real data is scarce, such as doctor–patient conversations.
I am currently a postdoctoral researcher in Prof. Shinji Watanabe's WAVLab at Carnegie Mellon University. I received my PhD in Information Engineering (2023) and MSc in Electronic Engineering (summa cum laude, 2019) from Università Politecnica delle Marche, advised by Prof. Stefano Squartini. Along the way I have worked with Amazon Alexa (applied scientist intern), consulted for startups building AI-powered hearing aids (Fortell, formerly Chromatic, which has since raised $163M at a $740M valuation) and silent speech interfaces (AlterEgo), and took part in three JHU JSALT workshops — as a participant in 2019 (during my master's) and 2026, and leading, with Lukáš Burget, the JSALT 2025 EMMA team on end-to-end multi-channel multi-talker ASR, which has produced six papers to date.
I believe strongly in open and reproducible science: I co-authored and contribute to several widely used open-source speech toolkits (SpeechBrain, ESPnet, Asteroid, TorchAudio — collectively exceeding 23,000 GitHub stars), and I organize community challenges — CHiME-7/8 DASR, URGENT, DCASE Task 4 — that push the field toward robustness and generalization. Core architectures I co-developed, such as SepFormer and TF-GridNet, are widely adopted by industrial research labs — including Google, Microsoft, MERL, and Samsung — for speech separation and enhancement. I am an elected member of the IEEE Speech and Language Processing Technical Committee (SLTC, Speech Recognition area, 2027–2029), served as secretary of the ISCA Special Interest Group on Robust Speech Processing (2023–2025), and routinely serve as meta-reviewer for ICASSP and Interspeech, with 60+ publications and 4,000+ citations in these areas.
When I am not doing research, I like to hike, ski (picture above), swim or — very recently — play tennis and disc golf (I am bad at both; I used to be decent at soccer though). I used to play the electric guitar; now I only play the keyboard (only the laptop one, sadly).
News
- Aug 2026Elected member of the IEEE Speech and Language Processing Technical Committee (SLTC), Speech Recognition area, for the 2027–2029 term.
- Aug 2026Took part in the JSALT 2026 workshop as a member of the Spoken Conversational AI Simulator team, on simulating full-duplex conversations for evaluating AI systems.
- Sep 2026Together with colleagues from NTT and Brno University of Technology, we will give a tutorial on Conversational Speech Recognition and Analysis: Progress, Challenges, and Emerging Opportunities with Speech Language Models at Interspeech 2026 in Sydney (28 September – 1 October 2026).
- Sep 2026Eight papers accepted at Interspeech 2026.
- Fall 2026Teaching CMU's 11-751/18-781 Speech Recognition and Understanding course, together with Jee-weon Jung.
- May 2026Invited talk at the JHU Human Language Technology Center of Excellence (HLTCOE).
- Feb 2026Invited talks at the CMU LTI Colloquium and the MILA Conversational AI reading group.
- Jan 2026MAPSS accepted at ICLR 2026; SE-DiCoW and the URGENT challenge accepted at ICASSP 2026. Invited talk at MBZUAI.
- Jan 2026Five papers from the JSALT 2025 EMMA team submitted to ICASSP 2026 — new state-of-the-art results on several multi-talker ASR benchmarks and a new multi-talker ASR framework based on shuffle automata.
- Dec 2025ASRU 2025 hackathon prize for best technical quality and correctness. ARECHO presented at NeurIPS 2025.
- Oct 2025Invited talks at Microsoft and the CMU Speech Lunch.
- Aug 2025Wrapped up the JSALT 2025 workshop in Brno, where I led with Lukáš Burget the 17-person EMMA team on end-to-end multi-channel multi-talker ASR — an effort that has since produced six papers.
- Jun 2025OpenBEATs (WASPAA 2025), ESPnet-SpeechLM (NAACL 2025), CS-FLEURS (Interspeech 2025) and the Interspeech 2025 URGENT challenge papers are out.
- 2024Awarded a Meta Audiobox Research Grant. Guest editor for the Computer Speech & Language special issue on multi-speaker, multi-microphone distant speech recognition.
Open-Source Software
- SpeechBrain — core developer and co-author of one of the most used PyTorch speech toolkits (10k+ stars).
- ESPnet — top-10 all-time contributor; co-author of ESPnet-SE++, ESPnet3 and the ESPnet-SpeechLM toolkit.
- Asteroid — co-author of the PyTorch audio source separation toolkit, natively integrated on the Hugging Face Hub (340+ models) and at one point featured on the Hugging Face landing page alongside organizations like fairseq and ESPnet.
- TorchAudio — co-author of TorchAudio 2.1.
- Contributor to Lhotse, DCASE-REPO, NVIDIA NeMo, Kaldi, and Pyroomacoustic.
Selected honors
- ASRU 2025 hackathon prize for best technical quality and correctness
- 1st place, 2nd Clarity Enhancement Challenge (2022) — demo
- 1st place, L3DAS22 Speech Enhancement Challenge (2022)
- 1st place (Developer Tools track), Global PyTorch Summer Hackathon 2020 — DeMask, built with Asteroid
- Judges' award for most innovative system, DCASE 2020 Task 4 Challenge
- Best student paper (co-author), IEEE SLT 2022
- Meta Audiobox Research Grant (2024)
- Outstanding Reviewer, IEEE ICASSP 2023
Press
- Quoted as an expert in MIT Technology Review: “A new AI translation system for headphones clones multiple voices simultaneously” (May 2025)
- Quoted as an expert in MIT Technology Review: “Noise-canceling headphones use AI to let a single voice through” (May 2024)
Teaching
- 11-751/18-781 Speech Recognition and Understanding — instructor (with Jee-weon Jung), Carnegie Mellon University, Fall 2026.
- Guest lectures and tutorials: JSALT 2025 summer school (robust speech recognition), CMU's Speech Technology for Conversational AI (2024, 2025), Université de Montréal, and UNIVPM.
Publications
Selected from 60+ publications — see Google Scholar for the full list.
Most cited
Recent
Contact
samuele.cornell [at] ieee.org
Language Technologies Institute, Carnegie Mellon University
5000 Forbes Avenue, Pittsburgh, PA
Always happy to chat about robust speech processing, open-source, challenges, or potential collaborations.