Change search
Link to record
Permanent link

Direct link
Publications (3 of 3) Show all publications
Arantes, P. & Eriksson, A. (2019). Quantifying Fundamental Frequency Modulation as a Function of Language, Speaking Style and Speaker. In: Interspeech 2019: . Paper presented at Interspeech 2019, Graz, Austria, September 15-19, 2019. Graz: The International Speech Communication Association (ISCA)
Open this publication in new window or tab >>Quantifying Fundamental Frequency Modulation as a Function of Language, Speaking Style and Speaker
2019 (English)In: Interspeech 2019, Graz: The International Speech Communication Association (ISCA), 2019Conference paper, Published paper (Refereed)
Abstract [en]

In this study, we outline a methodology to quantify the degree of similarity between pairs of f0 distributions based on the Anderson-Darling measure that underlies its namesake goodness-of-fit test. The procedure emphasizes differences due to more fine-grained f0 modulations rather than differences in measures of central tendency, such as the mean and median. In order to assess the procedure’s usefulness for speaker comparison, we applied it to a multilingual corpus in which participants contributed speech delivered in three speaking styles. The similarity measure was calculated separately as function of speaking style and speaker. Between-speaker variability (different speakers, same style) in distribution similarity varied significantly between styles — spontaneous interview shows greater variability than read sentences and word list in five languages (English, French, Italian, Portuguese and Swedish); in Estonian and German, read sentences yield more variability. Within-speaker variability (same speaker, different styles) levels are lower than between-speaker in the style that exhibit the greatest variability. The results point to the potential use of the proposed methodology as a way to identify possible idiosyncratic traits in f0 distributions. Also, they further demonstrate the effect of speaking styles on intonation patterns.

Place, publisher, year, edition, pages
Graz: The International Speech Communication Association (ISCA), 2019
Keywords
fundamental frequency, speaking style, cross- language comparison
National Category
General Language Studies and Linguistics
Research subject
Phonetics
Identifiers
urn:nbn:se:su:diva-195743 (URN)10.21437/Interspeech.2019-2857 (DOI)
Conference
Interspeech 2019, Graz, Austria, September 15-19, 2019
Funder
The Swedish Foundation for International Cooperation in Research and Higher Education (STINT), 1915339
Note

This work was jointly funded by Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES) grant 88881.155645/2017-01 as part of the STINT/CAPES research exchange agreement.

Available from: 2021-08-26 Created: 2021-08-26 Last updated: 2022-02-25Bibliographically approved
Arantes, P., Eriksson, A. & Lima, V. (2018). Minimum Sample Length for the Estimation of Long-term Speaking Rate. Paper presented at 9th International Conference on Speech Prosody 2018, Poznań, Poland, June 13-16, 2018. Speech prosody, 661-665
Open this publication in new window or tab >>Minimum Sample Length for the Estimation of Long-term Speaking Rate
2018 (English)In: Speech prosody, ISSN 2333-2042, p. 661-665Article in journal (Refereed) Published
Abstract [en]

In this study, we expand on previous experiments designed with the aim of determining the minimum length that an audio sample should have in order for the speaking rate derived from it to be representative of the sample as a whole. We compare two different approaches to establishing that the time series of the cumulative speaking rate calculated over the audio sample has reached stability. We also compare the effect on stabilization time of four other factors that may affect the way speaking rate is calculated. The results show that all factors tested have significant effects, although of limited practical concern. Overall, average stability time is 12.1 seconds, with the bulk of the distribution lying between 7.9 and 16.2 s.

Keywords
speaking rate, speech rate, articulation rate, forensic phonetics
National Category
General Language Studies and Linguistics
Research subject
Arts, Humanities and Social Science Education; Phonetics
Identifiers
urn:nbn:se:su:diva-184060 (URN)10.21437/SpeechProsody.2018-134 (DOI)
Conference
9th International Conference on Speech Prosody 2018, Poznań, Poland, June 13-16, 2018
Note

This work has been supported by STINT (Sweden) grantIB2015-6488. The third author was supported by FAPESP(Brazil) grant 2016/12646-0 between November 2016 and December2017.

Available from: 2021-02-20 Created: 2021-02-20 Last updated: 2022-02-25Bibliographically approved
Arantes, P., Eriksson, A. & Gutzeit, S. (2017). Effect of Language, Speaking Style and Speaker on Long-Term F0 Estimation. Paper presented at INTERSPEECH 2017, Stockholm, Sweden, August 20–24, 2017. Interspeech, 3897-3901
Open this publication in new window or tab >>Effect of Language, Speaking Style and Speaker on Long-Term F0 Estimation
2017 (English)In: Interspeech, ISSN 2308-457X, p. 3897-3901Article in journal (Refereed) Published
Abstract [en]

In this study, we compared three long-term fundamental frequency estimates — mean, median and base value — with respect to how fast they approach a stable value, as a function of language, speaking style and speaker. The base value concept was developed in search for an f0 value which should be invariant under prosodic variation. It has since also been tested in forensic phonetics as a possible speaker-specific f0 value. Data used in this study — recorded speech by male and female speakers in seven languages and three speaking styles, spontaneous, phrase reading and word list reading — had been recorded for a previous project. Average stabilisation times for the mean, median and base value are 9.76, 9.67 and 8.01 s. Base values stabilise significantly faster. Languages differ in both average and variability of the stabilisation times. Values range from 7.14 to 11.41 (mean), 7.5 to 11.33 (median) and 6.74 to 9.34 (base value). Spontaneous speech yields the most variable stabilisation times for the three estimators in Italian and Swedish, for the median in French and Portuguese and base value in German. Speakers within each language do not differ significantly in terms of stabilisation time variability for the three estimators.

 

Keywords
acoustic phonetics, speech acoustics, forensic phonetics
National Category
General Language Studies and Linguistics
Research subject
Phonetics
Identifiers
urn:nbn:se:su:diva-190500 (URN)10.21437/Interspeech.2017-449 (DOI)
Conference
INTERSPEECH 2017, Stockholm, Sweden, August 20–24, 2017
Note

This work has been supported by a grant (IB2015-6488) from The Swedish Foundation for International Cooperation in Research and Higher Education (STINT). The third author was supported by a PIBIC grant by Conselho Nacional de Desenvolvimento Científico e Tecnológico(CNPq) between August 2014 and July 2015

Available from: 2021-02-20 Created: 2021-02-20 Last updated: 2022-02-25Bibliographically approved
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-9707-8493

Search in DiVA

Show all publications