| Abstract | With advances in AI, audio deepfakes are getting more and more convincing. Audio deepfakes can be generated using Text-to-Speech synthesis or Voice conversion (VC). The evaluation of these samples is still an ongoing research domain. This paper analyses 77 VC papers published in ICASSP and Interspeech in 2023 and 2024, focusing on the use of the Mean Opinion Score (MOS) which is often used for evaluating the naturalness, intelligibility, or quality of generated speech samples. First, we examined how MOS tests are described and applied in scientific papers. Second, we analysed the reported MOS across different papers to assess their reusability and interpretability. The first analysis reveals the general trends in the current use of MOS tests. For instance, we observed that certain aspects, such as the number of raters, are reported in several papers. However, there is a noticeable lack of detailed information on the audio samples used in MOS tests. The second analysis highlights the need for more information regarding the generation of audio samples using baseline models. Consequently, we found that VC papers often face challenges in interpreting MOS and ensuring reproducibility, which should be considered when discussing and interpreting the MOS results. |
|---|