Neural Audio Codecs
Discrete neural audio codecs for speech compression and autoregressive modeling — DAC, SNAC, WavTokenizer, and X-Codec2.
Build & Train Models
Training pipelines, tokenizers, and infrastructure for large-scale audio and language models.
Speech Alignment
Extracting precise word and phoneme timestamps from audio — tools, models, and common failure modes.
Speech Datasets
Clean speech corpora, noise and RIR databases, and techniques for synthesizing degraded training data.
Voice Conversion
Transforming a speaker's voice to match a target speaker while preserving linguistic content.
Speech Editing
Modifying spoken audio at the word level — insertions, deletions, and substitutions while preserving speaker identity.
Speech Enhancement
Reconstructing clean speech from degraded audio — models, metrics, and inference setups.