Enhancing Neural Vocoders with Fourier Transform: A Frequency-Domain Approach to Improved Speech Synthesis
DOI:
https://doi.org/10.61173/ggbyhg71Keywords:
Neural vocoders, Fourier Transform, Speech synthesis, Noise reductionAbstract
This paper introduces a frequency-domain approach to enhance neural vocoders, addressing limitations in capturing high-frequency details essential for natural and clear speech synthesis. By integrating Short-Time Fourier Transform preprocessing, the method provides two key benefits. It offers a richer, frequency-detailed input, enabling the vocoder to capture finer spectral elements for improved synthesis quality. It also facilitates targeted noise reduction, refining output clarity. Additionally, frequency-domain enhancements during generation allow selective amplification of key frequencies (e.g., 3–6 kHz). The integration of these techniques not only significantly improves the clarity and naturalness of synthesized speech but also reduces artifacts that commonly affect high-frequency content. By leveraging both traditional signal processing methods and deep learning, this framework enhances the vocoder‘s ability to accurately reproduce challenging speech spectra, providing a balanced approach that can generalize across different acoustic environments. Combining Fourier-based processing with neural networks, this approach pushes the boundaries of vocoder quality, improving both naturalness and intelligibility. These advancements set a new standard in speech synthesis, offering broader applications in audio processing.
References
[1] Kong, J., Kim, J., & Bae, J. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. Advances in Neural Information Processing Systems, 2020, 33, 17022-17033.
[2] Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde- Farley, D., Ozair, S., Courville, A., & Bengio, Y. Generative Adversarial Networks. Advances in Neural Information Processing Systems (NIPS), 2014, 2672-2680.
[3] Hinton, G., Vinyals, O., & Dean, J. Distilling the Knowledge in a Neural Network. arXiv preprint arXiv:1503.02531.
[4] Wang, Y., Stanton, D., Zhang, Y., Skerry-Ryan, R., Battenberg, E., Shor, J., Xiao, Y., Ren, F., Jia, Y., & Saurous, R. A. Tacotron: Towards End-to-End Speech Synthesis. Proceedings of Interspeech, 2017, 4006-4010.
[5] Siuzdak, H. Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis. arXiv preprint arXiv:2306.00814.
[6] Kong, J., Kim, J., & Bae, J. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. Advances in Neural Information Processing Systems, 2020, 33: 17022-17033.
[7] Prenger, R., Valle, R., & Catanzaro, B. WaveGlow: A Flow-based Generative Network for Speech Synthesis. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019, 3617-3621.
[8] Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., & Kavukcuoglu, K. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499.
[9] F Prenger, R., Valle, R., & Catanzaro, B. WaveGlow: A Flow-based Generative Network for Speech Synthesis. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019, 3617-3621.
[10] Wang, Y., Stanton, D., Zhang, Y., Skerry-Ryan, R., Battenberg, E., Shor, J., Xiao, Y., Ren, F., Jia, Y., & Saurous, R. A. Tacotron: Towards End-to-End Speech Synthesis. Proceedings of Interspeech, 2017, 4006-4010.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
