Speaker Diarization

term_id: speaker_diarization

Category: basic_concepts

Definition

Speaker Diarization is the task of partitioning an audio stream into homogeneous segments according to the identity of the speaker. It combines speaker change detection with speaker clustering to label segments with unique speaker IDs. This technology is essential for making multi-party conversations understandable in transcripts, often referred to as the ‘who said what’ problem.

Summary

The process of determining ‘who spoke when’ in an audio recording.

Key Concepts

  • Speaker clustering
  • Identity labeling
  • Who-said-what
  • Audio segmentation

Use Cases

  • Automatic meeting minutes generation
  • Interview transcription
  • Broadcast media analysis