Multimodal Speaker Verification as a Threat to Speaker Anonymization
Abstract
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.
Community
Most automatic speaker verification (ASV) systems
operate on individual utterances, despite real-world interactions
typically consisting of multiple utterances. As speech accumu-
lates, increasingly rich speaker information becomes available
through acoustic, prosodic, and linguistic cues, potentially chal-
lenging speaker anonymization methods that primarily target
vocal characteristics. We investigate ASV in a multi-utterance,
multimodal setting and examine whether aggregating information
across anonymized speech impacts privacy. We first study audio-
only aggregation across multiple anonymized utterances and
observe consistent performance improvements as more speech
becomes available. We then incorporate prosodic and linguistic
information, showing that multimodal systems outperform uni-
modal approaches. Finally, we compare aggregation strategies
and find that frame-level aggregation yields the lowest EERs.
Even with only five anonymized utterances, combining audio
and text reduces EER by over 15% relative to audio-only ag-
gregation, demonstrating that substantial speaker-discriminative
information remains accessible despite anonymization.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models (2026)
- Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment (2026)
- Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models (2026)
- NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization (2026)
- A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization (2026)
- Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification (2026)
- ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.19636 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper