
WAND: Efficient Autoregressive TTS with Windowed Attention and Knowledge Distillation
Researchers introduce WAND, a framework that reduces the computational and memory costs of autoregressive text-to-speech models by using windowed attention and knowledge distillation. This approach maintains high-fidelity speech output while making the models more scalable.
