Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Anonymous speaker clusters: Making distinctions between anonymised speech recordings with clustering interface

Domain:

natural language processing

Record type:

paper
Creator:
O'BTomChaBon
Editor:
LabLabANRANR
Publisher:
CCSD
Host:avatar
International audience Our study examined the performance of evaluators tasked to group natural and anonymised speech recordings into clusters based on their perceived similarities. Speech stimuli were selected from the VCTK corpus; two systems developed for the VoicePrivacy 2020 Challenge were used for anonymisation. The Baseline-1 (B1) system was developed by using x-vectors and neural waveform models, while the Baseline-2 (B2) system relied on digital-signal-processing techniques. 74 evaluators completed three trials composed of 16 recordings with either natural or anonymised speech generated from a single system. F-measure and cluster purity metrics were used to assess evaluator accuracy. Probabilistic linear discriminant analysis (PLDA) scores from an automatic speaker verification system were generated to quantify similarity between recordings and used to correlate subjective results. Our findings showed that non-native English speaking evaluators significantly lowered their F-measure means when presented anonymised recordings. We observed no significance for cluster purity. Pearson correlation procedures revealed that PLDA scores generated from natural and B2-anonymised speech recordings correlated positively to F-measure and cluster purity metrics. These findings show evaluators were able to use the interface to cluster natural and anonymised speech recordings and suggest anonymisation systems modelled like B1 are more effective at suppressing identifiable speech characteristics.

Visit

doi.org

Tasks

speaker verificationspeech processing

Tags

privacyanonymisationspeech synthesisspeaker identificationclusteringsubjective evaluation[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][SCCO.LING]Cognitive science/Linguistics

Licenses

info:eu-repo/semantics/OpenAccess

Similar

Yoruba Multi-Speaker Speech CorpusDo Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?Speech transcription platform user interfaceDendi of Parakou multi-speaker speech datasetTowards a speech recognition based automatic telephone exchange with an Afrikaans conversational interfaceA speaker independent continuous speech recognizer for Amharic

Yoruba Multi-Speaker Speech Corpus

Yoruba tts notebook and data

Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?

Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models,

Speech transcription platform user interface

This is the user interface component of the Speech Transcription Platform developed by the Multiling

Dendi of Parakou multi-speaker speech dataset

This dataset was created for speech research purposes and contains about 676 recordings of participa

Towards a speech recognition based automatic telephone exchange with an Afrikaans conversational interface

A speaker independent continuous speech recognizer for Amharic