2026
WorldSpeech: A Multilingual Speech Corpus from Around the World
Under Review
@misc{asonitis2026worldspeechmultilingualspeechcorpus,
title={WorldSpeech: A Multilingual Speech Corpus from Around the World},
author={Antonis Asonitis and Luca A. Lanzendörfer and Frédéric Berdoz and Roger Wattenhofer},
year={2026},
eprint={2605.09167},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.09167},
}
65k hours across 80+ languages, sourced from government proceedings, international broadcasts and public-domain audiobooks. 150k+ HuggingFace downloads; #1 trending in its first month of release.
High-Fidelity Speech Enhancement via Discrete Audio Tokens
ICASSP 2026
@InProceedings{Lanzendörfer2026highfidelity,
author= {Luca Lanzendörfer and Frédéric Berdoz and Antonis Asonitis and Roger Wattenhofer},
title= {{High-Fidelity Speech Enhancement via Discrete Audio Tokens}},
booktitle= {{Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Catalonia, Spain}},
month= {May},
year= {2026},
}
Post-Training Speech Enhancement Language Models with Perceptual Rewards
Oral, Interspeech 2026
BibTeX
arXiv
@misc{berdoz2026posttrainingspeechenhancementlanguage,
title={Post-Training Speech Enhancement Language Models with Perceptual Rewards},
author={Frédéric Berdoz and Luca A. Lanzendörfer and Antonis Asonitis and Roger Wattenhofer},
year={2026},
eprint={2606.21458},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2606.21458},
}
Multilingual Speech Editing
Oral, ICML 2026 Workshop on Machine Learning for Audio
BibTeX
Samples
@misc{asonitis2026multilingualspeechediting,
title = {Multilingual Speech Editing},
author = {Antonis Asonitis and Luca A. Lanzend{\"o}rfer and Fr{\'e}d{\'e}ric Berdoz and Roger Wattenhofer},
year = {2026},
note = {Under review}
}
Autoregressive Zero-Shot Voice Conversion
Oral, ICML 2026 Workshop on Machine Learning for Audio
Semantic Calibration in Media Streams
Under Review
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Under Review
The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
Under Review
Controlling Speaking Rate in Autoregressive TTS via Activation Steering
Under Review
OmniTag-LLM: A Multimodal Large Language Model for Multilingual Multitask Speech Labeling
Under Review
Datasets
Multilingual Audio Alignments
The largest multilingual word and phoneme aligned speech dataset, 45k+ downloads, reached #3 Trending Worldwide in the Audio domain.