DER (Diarization Error Rate)
Diarization error rate is the standard measure of who-spoke-when accuracy. It sums three kinds of mistake — speech attributed to the wrong speaker, speech the system missed, and speech it detected where there was none — and divides the total by the amount of speech in the reference annotation. Lower is better.
Two scoring choices change the number substantially. A forgiveness collar excludes a short window on either side of each speaker change, where annotation boundaries are least reliable; a 250 ms collar is common. Evaluations also differ in whether overlapping speech is included or excluded. Both choices remove the hardest moments from the calculation, so DER figures are only comparable when the collar and the overlap handling are stated alongside them.