WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability

Yongkang Yin; Yuexian Zou

doi:10.1609/aaai.v40i41.40746

Authors

Yongkang Yin Guangdong Provincial Key Laboratory of Ultra High Definition lmmersive Media Technology, Shenzhen Graduate School, Peking University ADSPLAB, School of ECE, Peking University, Shenzhen, China
Yuexian Zou Guangdong Provincial Key Laboratory of Ultra High Definition lmmersive Media Technology, Shenzhen Graduate School, Peking University ADSPLAB, School of ECE, Peking University, Shenzhen, China

DOI:

https://doi.org/10.1609/aaai.v40i41.40746

Abstract

Speaker diarization is a fundamental task in speech processing aims to determine 'who speaks when'. When combined with ASR, it enables speaker-labeled transcription with broad practical value. Most existing methods rely on frame-level classification, but the high cost of annotating mixed-speaker audio limits the availability of large-scale, accurately labeled datasets. As a result, even state-of-the-art models struggle with imprecise speaker boundary detection and semantic segmentation errors, which degrade timestamp accuracy and downstream ASR performance. To address these challenges, we propose WhisperDiari, a unified framework for speaker diarization and ASR. We first construct LibriDiari, a dataset derived from LibriSpeech, containing 2–4 speaker mixed audio annotated with transcripts and speaker labels. WhisperDiari builds on the Whisper model, incorporating speaker adapters and Speaker Similarity Matrix Supervision to enhance speaker representation. In addition, a dedicated speaker decoder fuses speaker embeddings with contextual semantics from Whisper's decoder, enabling token-level diarization. This design effectively resolves segmentation ambiguity, aligns diarization with semantic units, and jointly models 'who speaks what and when', producing accurate, timestamped transcripts. We train the model on LibriDiari and evaluate it on both LibriDiari and the real-world AMI corpus. Experimental results demonstrate that WhisperDiari consistently outperforms state-of-the-art open-source baselines.

WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information