Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Abstract
Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate persistent speaker attribution with Speaker Identified cpWER (SI-cpWER), which scores a corpus under one global speaker-ID assignment. The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6. ThyVoice is our end-to-end reference system; it repairs overlap and gates the evidence used to create and update voiceprints. Requiring persistent identity changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions and the lowest mean in the full panel, 47.13 versus 54.75 for the next system. Complementary lexical, diarization, per-recording attribution, and speaker-clustering diagnostics characterize upstream error surfaces in the final attributed record. These results show why persistent attribution must be evaluated directly in systems that reuse conversations across time.
Community
Meeting transcripts are usually scored one recording at a time, so a system can look strong while assigning the same person different identities across meetings. We measure this with SI-cpWER: cpWER under one speaker mapping across the whole corpus.
We evaluate 5 commercial diarized-STT APIs and 2 open baselines, each paired with the same pyannoteAI voiceprint backend, plus our reference system ThyVoice, on NOTSOFAR (clean and noise-augmented) and CHiME-6. The commercial ranking changes: the best per-meeting cpWER does not give the best cross-meeting attribution.
The report also breaks errors down by layer (WER, DER/JER, speaker clustering) and tests how overlap handling and enrollment audio affect identity.
A short version is accepted at IEEE SLT 2026 (Demo Track).
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- FuseAlign: Forced Alignment in the Wild (2026)
- STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR (2026)
- Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations (2026)
- VibeVoice-ASR-Streaming Technical Report (2026)
- FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy Conditions (2026)
- A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations (2026)
- TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper