Percorrer por autor "Casanova, Edresson"
A mostrar 1 - 2 de 2
Resultados por página
Opções de ordenação
- Interpretability analysis of deep models for COVID-19 detectionPublication . Silva, Daniel; Casanova, Edresson; Gris, Lucas; Gauy, Gustavo; Junior, Arnaldo; Finger, Marcelo; Svartman, Flaviane; Medeiros, Beatriz; Martins, Marcus; Aluísio, Sandra; Berti, Larissa; Teixeira, João PauloDuring the coronavirus disease 2019 (COVID-19) pandemic, various research disciplines collaborated to address the impacts of severe acute respiratory syndrome coronavirus-2 infections. This paper presents an interpretability analysis of a convolutional neural network-based model designed for COVID-19 detection using audio data. We explore the input features that play a crucial role in the model’s decision-making process, including spectrograms, fundamental frequency (F0), F0 standard deviation, sex, and age. Subsequently, we examine the model’s decision patterns by generating heat maps to visualize its focus during the decision-making process. Emphasizing an explainable artificial intelligence approach, our findings demonstrate that the examined models can make unbiased decisions even in the presence of noise in training set audios, provided appropriate preprocessing steps are undertaken. Our top-performing model achieves a detection accuracy of 94.44%. Our analysis indicates that the analyzed models prioritize high-energy areas in spectrograms during the decision process, particularly focusing on high-energy regions associated with prosodic domains, while also effectively utilizing F0 for COVID-19 detection.
- TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian PortuguesePublication . Casanova, Edresson; Junior, Arnaldo Candido; Shulby, Christopher; Oliveira, Frederico Santos de; Teixeira, João Paulo; Ponti, Moacir Antonelli; Aluísio, SandraSpeech provides a natural way for human–computer interaction. In particular, speech synthesis systems are popular in different applications, such as personal assistants, GPS applications, screen readers and accessibility tools. However, not all languages are on the same level when in terms of resources and systems for speech synthesis. This work consists of creating publicly available resources for Brazilian Portuguese in the form of a novel dataset along with deep learning models for end-to-end speech synthesis. Such dataset has 10.5 h from a single speaker, from which a Tacotron 2 model with the RTISI-LA vocoder presented the best performance, achieving a 4.03 MOS value. The obtained results are comparable to related works covering English language and the state-of-the-art in European Portuguese.
