Real-Time Speaker Diarization

Real-Time Speaker Diarization

●2 ●9
calendar_today ago • schedule12 min read

Real-Time Speaker Diarization

Algorithms, web implementation and applications in drone detection by the acoustic fingerprint of their motors

This article presents the principles and practical implementation of speaker diarization – the automatic identification of the person speaking at a given moment in a continuous audio stream. Starting from the Fourier transform, the Mel scale and the MFCC coefficients, we explore classical and modern speaker identification algorithms, as well as the challenges of implementing them in the browser using the Web Audio API. In the second part, we show that the same set of techniques, applied to the signals emitted by the electric motors of unmanned aerial vehicles (UAVs/drones), can form the basis of an acoustic drone detection and identification system – a field with applications in security, air traffic management and critical infrastructure protection.

1. INTRODUCTION

In any conversation involving several participants – from a medical consultation to a board meeting or a journalistic interview – the question “who spoke?” is just as important as “what was said?”. Speaker diarization is the task of segmenting an audio stream into homogeneous regions attributed to a single speaker, answering the question “who speaks when?” without assuming prior knowledge of the identity of the people involved [1].

The field has evolved dramatically over the last decade: while in the 2000s systems relied on Gaussian mixture models (GMM) and the BIC distance, today deep neural architectures produce speaker embeddings (x-vectors, d-vectors) with performance exceeding human discrimination ability under noisy conditions [2, 3]. Nevertheless, efficient implementation in resource-constrained environments – browser, embedded systems, edge computing – remains an open challenge of immediate practical interest.

This article has a twofold purpose:

  1. to present, in sufficient technical detail, the mathematical and algorithmic foundations of speaker diarization, with an emphasis on implementability in JavaScript using the Web Audio API; and

  2. to show that the same acoustic feature extraction pipeline can be adapted for detecting and identifying drones by the sound signature of their motors – an active research field with applications in national security, urban planning and airspace protection [4, 5].

  3. FROM WAVEFORM TO FINGERPRINT: AUDIO SIGNAL PROCESSING

2.1 The Fourier transform and the frequency domain

The raw audio signal is a sequence of acoustic pressure values sampled at a fixed rate (typically 44,100 Hz for consumer audio). Although the time-domain waveform contains all the information, it is not suitable for extracting discriminative speaker features. The Discrete Fourier Transform (DFT) – and its efficient implementation, the FFT – projects the signal into the frequency domain, revealing the distribution of energy across sinusoidal components.

In practice, the analysis is performed on short windows (typically 20–40 ms), with overlap (10–15 ms), yielding a time-frequency representation called a spectrogram. Each window is multiplied by a window function (Hamming, Hann) to reduce the effect of spectral leakage at the window edges.

2.2 The Mel scale and the filterbank

Human perception of pitch is not linear with respect to physical frequency: the ear discriminates low frequencies more finely than high ones. The Mel scale, proposed by Stevens, Volkmann and Newman in 1937, approximates this non-linearity:

A mel filterbank consists of M triangular filters distributed uniformly on the Mel scale, applied to the FFT power spectrum. The energy of each filter is log-compressed (to compress the dynamic range, similarly to the auditory system), yielding a vector of M values representing the mel spectrum of the current window. Typical values are M = 26 or 40.

2.3 Mel-Frequency Cepstral Coefficients (MFCC)

The MFCC coefficients, introduced by Davis and Mermelstein in 1980 [6], are obtained by applying the Discrete Cosine Transform (DCT) to the log-energies of the mel filterbank. The DCT decorrelates the components (since the mel filterbanks overlap), condensing the vocal information into the first 12–13 coefficients:

The coefficient c₀ (k=0) corresponds to the total energy of the frame and is often excluded from the feature vector for speaker identification, since it reflects the loudness of speech rather than the identity of the speaker. The MFCC vectors c₁–c₁₃ capture the characteristics of the vocal tract – the shape of the oral cavity, its length, the position of the tongue – stable features of vocal identity.

The complete feature vector for a frame often also includes the delta coefficients (Δ-MFCC, the temporal derivative) and delta-delta coefficients (ΔΔ-MFCC), which capture the temporal dynamics of speech. To these are sometimes added the fundamental frequency (pitch, F₀) – estimated by autocorrelation on the time-domain signal – and the spectral centroid, which measures the “brightness” of the voice.

Key point: The performance of any voice identification system depends critically on the quality of feature extraction. Cepstral Mean Normalization (CMN) – subtracting the session mean from each MFCC vector – reduces the influence of channel characteristics (microphone, room acoustics). L2 normalization (projection onto the unit sphere) is a robust, temporally stable alternative that allows the Euclidean distance to be used as a consistent comparison metric.

3. SPEAKER IDENTIFICATION AND DIARIZATION ALGORITHMS

3.1 Gaussian Mixture Models (GMM-UBM)

The classical approach, dominant in the 2000s, models the distribution of each speaker’s MFCC vectors with a Gaussian Mixture Model (GMM). A Universal Background Model (UBM) trained on a large corpus of speakers is adapted by Maximum A Posteriori (MAP) estimation to the data of each specific speaker – yielding a personalized GMM at low computational cost [7]. Identification is performed by a likelihood ratio test against the UBM.

The limitations of GMM-UBM are its sensitivity to channel variations and the need for at least a few tens of seconds of training data per speaker. These limitations motivated the development of methods based on latent subspaces.

3.2 I-vectors and Factor Analysis

Dehak et al. (2011) [2] proposed representing a speaker session as a low-dimensional vector (typically 400) – the i-vector – which simultaneously encodes speaker identity and channel variables in a total variability subspace. The model assumes that the GMM supervector can be written as the sum of a background supervector and a total variability term: , where T is the total variability matrix and w is the session’s i-vector.

Classification is performed with Linear Discriminant Analysis (LDA) or Probabilistic LDA (PLDA), which separate intra-speaker variability from inter-speaker variability. I-vectors dominated the NIST Speaker Recognition Evaluation (SRE) competitions for almost a decade.

3.3 X-vectors and Deep Neural Networks

Deep neural architectures have redefined the state of the art in speaker identification. Snyder et al. (2018) [3] propose x-vectors: a DNN model trained for speaker classification on windows of a few seconds, from whose intermediate layers compact embeddings (typically 512-dimensional) are extracted. A statistical aggregation layer (mean + standard deviation pooling) allows processing of variable-length sequences.

X-vectors, combined with PLDA and data augmentation (added noise, simulated reverberation), reduced the identification error rate by 50–70% compared to i-vectors on standard benchmarks. More recent models – ECAPA-TDNN, ResNet-based speaker models – achieve even higher performance, with an EER (Equal Error Rate) below 1% on clean datasets [1, 8].

3.4 Diarization vs. Verification vs. Identification

Three related tasks must be distinguished: (1) Speaker verification – is person X the one speaking? (binary classification); (2) Speaker identification – which of N known speakers is speaking? (multi-class classification); (3) Speaker diarization – segmenting the audio stream into regions attributed to speakers without prior knowledge of their identity, with the number of speakers unknown a priori [1].

A typical diarization pipeline includes:

  1. voice activity detection (VAD),

  2. feature extraction,

  3. clustering (k-means, spectral clustering, or DBSCAN on embeddings),

  4. optionally re-segmentation with HMM or Viterbi.

The standard evaluation metric is the DER (Diarization Error Rate), the sum of the missed speech, false alarm and speaker confusion rates.

4. BROWSER IMPLEMENTATION: WEB AUDIO API AND NEAREST-CENTROID

4.1 The Web Audio API and the constraints of the browser environment

The Web Audio API [9] natively provides, in the browser, access to the microphone stream via getUserMedia(), signal processing with AnalyserNode (real-time FFT, configured via fftSize = 2048), and precise timing via AudioContext. Implementation in JavaScript imposes significant constraints compared to server-side systems: the lack of explicit SIMD for matrix operations, single-threaded execution on the main thread (although Web Workers and AudioWorklets partially alleviate this), and a latency of ~16 ms imposed by requestAnimationFrame.

These constraints make it impossible to run large DNN models (x-vectors, ECAPA-TDNN) directly in the browser without ONNX Runtime Web or TensorFlow.js – and even with these frameworks, real-time inference on models with millions of parameters is only marginally feasible on modest hardware. The practical alternative is a lightweight classifier: nearest-centroid with a Euclidean metric in the L2-normalized feature space.

4.2 The real-time feature extraction pipeline

The practical implementation processes frames of 2048 samples (≈46 ms at 44.1 kHz) on each requestAnimationFrame cycle. Only frames with RMS > θ (a configurable energy threshold) are included in the collection buffer of the current utterance – a simplified voice activity detection that is nevertheless effective under low-noise conditions.

At the end of an utterance (detected when speech recognition stops or via a silence timer), the raw MFCC vectors of the collected frames are averaged, and the mean is L2-normalized – yielding the L2 fingerprint of the current utterance:

This order – averaging the raw vectors, followed by normalization – is critical. Per-frame normalization followed by averaging causes vector cancellation: the mean of vectors on the unit sphere may be a vector of near-zero norm, devoid of discriminative power. Normalizing the raw mean produces a stable centroid, reproducible for the same speaker.

4.3 The nearest-centroid classifier with adaptive threshold

Each detected speaker is associated with a centroid – the L2-normalized (and re-normalized) mean of the last K fingerprints. A new utterance is identified by its Euclidean distance to each known centroid:

If the minimum distance exceeds a threshold θ, a new speaker is declared. The optimal threshold is , where is the Euclidean distance between the two known centroids – this is calibrated automatically after the first two speakers are detected. Before calibration, the initial threshold θ₀ = 0.5 works reasonably in the L2-normalized space, where typical intra-speaker distances are 0.05–0.25 and inter-speaker distances 0.5–1.4.

A further improvement: after K₀ = 5 utterances from the first speaker, the threshold is recalibrated from the observed intra-speaker variance: , where is the mean Euclidean distance of the samples from their centroid. This makes the system adaptive to the type of voice and the acoustic conditions of the room.

5. DRONE DETECTION BY ACOUSTIC FINGERPRINT

5.1 The acoustic signature of UAVs

Multirotor unmanned aerial vehicles (quadcopters, hexacopters, etc.) produce a characteristic sound signal, determined by three main sources:

  1. the rotation of the brushless electric motors, which generates tones at the switching frequency of the ESC controllers and their harmonics;

  2. the aerodynamic sound of the propellers – air-cutting noise (blade passing frequency, ) and blade-tip turbulence;

  3. the structural resonance of the frame.

The BPF is the dominant and distinctive frequency: a typical quadcopter (DJI Phantom 4) with 2-blade propellers and RPM ≈ 8,000–10,000 produces a BPF ≈ 267–333 Hz. Models with different configurations (hex, octo, multiple blades) have distinct, identifiable spectral fingerprints [4, 5].

Unlike the human voice – wide band 80–8,000 Hz, variable formant structure – the sound of a drone motor is quasi-periodic, with energy concentrated in the BPF and its harmonics (2×BPF, 3×BPF...). This makes identification easier under ideal conditions, but more fragile to RPM variations (which depend on the maneuver being executed) and to urban environmental noise.

5.2 Transferring voice identification techniques to UAVs

Jeon et al. (2017) [4] demonstrated that a deep neural network trained on mel spectrograms of drone sounds can detect the presence of a drone with >95% accuracy at distances of up to 80 m in semi-controlled environments. Al-Emadi et al. (2019) [5] extended the approach to identifying the specific model (DJI Phantom vs. Parrot Bebop vs. background), using MFCC and CNN, obtaining 98% accuracy on the test set.

The parallel with speaker identification is direct: instead of the vocal tract, the acoustic signature is determined by the motor + propeller configuration; instead of formants, we track the BPF and its harmonics. The technical pipeline:

FFT → mel filterbank → MFCC → L2 normalization → nearest-centroid or DNN classifier

transfers almost without modification, with a few domain-specific adjustments:

  • Frequency band of interest: 100–4,000 Hz (vs. 80–8,000 Hz for voice);

  • FFT window size: 4,096–8,192 samples for better frequency resolution at low BPFs;

  • Number of mel filters: 32–64 (more, in order to capture the fine harmonics);

  • Augmentation with synthetic pitch-shifting to cover RPM variations.

5.3 Multi-microphone acoustic detection and localization

One advantage of acoustic detection over radar or LiDAR is the low cost of the sensors. A microphone array allows not only detecting presence, but also estimating the direction of arrival (DOA) through beamforming or MUSIC (Multiple Signal Classification). The time difference of arrival (TDOA) between pairs of microphones, combined with triangulation, can localize the drone in 3D space.

Emerging commercial systems (DroneShield, Dedrone – although predominantly based on RF and video) integrate acoustic detection as an additional layer, especially for small drones with a low radar signature. Recent research explores distributed microphone networks, processed at the edge computing level with centralized aggregation – an architecture that benefits directly from the lightweight feature extraction techniques discussed in Section 4.

5.4 Challenges and limitations

The main challenges of acoustic drone detection are:

  1. Urban noise – road traffic, air conditioning, wind – spectrally overlapping with the BPF frequencies;

  2. RPM variation depending on the maneuvers executed, which shifts the BPF by ±30%;

  3. The maximum effective range, limited to 50–100 m in urban conditions (vs. 200–300 m in rural environments);

  4. Detection latency – a rapid-response system requires inference under 200 ms, which limits model complexity.

The same challenges exist, by analogy, in voice identification in noisy environments: overlapping speech, variable background noise, emotional variations in vocal pitch. The solutions developed in ASR and speaker recognition – noise data augmentation, adaptive beamforming, robust VAD based on spectral energy – are directly applicable to acoustic UAV detection.

6. COMPARATIVE EVALUATION AND FUTURE DIRECTIONS

6.1 Lightweight vs. DNN: when is nearest-centroid sufficient?

The nearest-centroid classifier in the L2-MFCC space is viable in scenarios with a small number of speakers (2–4), relatively controlled acoustic conditions and latency requirements under 50 ms (execution in the browser or on ARM Cortex-M4+ microcontrollers). On the VoxCeleb1 dataset (7,000+ speakers), nearest-centroid with MFCC yields an EER of approximately 15–20%, compared to 1–3% for DNN models. The gap narrows significantly when the number of speakers is small and training data is abundant [1, 3].

For drone detection, where the number of models to identify is small (tens, not thousands) and the acoustic signature is more stable than the human voice, nearest-centroid MFCC achieves performance comparable to CNN models on benchmarks with SNR > 10 dB [4].

6.2 Integrating pre-trained models via ONNX Runtime Web

A promising direction for improving browser performance without sacrificing accessibility is the use of ONNX Runtime Web with WebAssembly (and, where available, WebGPU). Small models of the SpeakerNet or audio MobileNet type, exported in ONNX format, can run inference for a 2-second frame in under 100 ms on mainstream hardware, paving the way for lightweight x-vectors directly in the browser.

Similarly, pre-trained drone detection models (CNNs on 1–2 second mel spectrograms, with ~500K parameters) are small enough for ONNX Web deployment, with potential use in security applications accessible to the general public (e.g. mobile apps for citizens reporting aerial intrusions).

6.3 Privacy and ethical considerations

Voice identification systems raise legitimate privacy questions: collecting voiceprints constitutes processing of biometric data under the GDPR and similar regulations. Implementations that process exclusively locally (browser, edge) and do not transmit audio data or fingerprints to a server offer stronger guarantees than cloud solutions. The privacy-by-design principle recommends: local processing, volatile fingerprints (reset when the session closes), and explicit user consent.

In the drone domain, acoustic detection does not raise the same privacy issues as video surveillance, but national regulations on the interception of communications and airspace monitoring require clear legal frameworks for the operators of detection systems.

7. CONCLUSIONS

Speaker diarization – the automatic identification of the person speaking in an audio stream – relies on an elegant mathematical pipeline:

FFT → mel filterbank → MFCC → L2 normalization → nearest-centroid or DNN classification.

Browser implementation with the Web Audio API is feasible for scenarios with two to four speakers, provided that a few critical details are respected: excluding c₀ from the MFCC vector, averaging the raw vectors before normalization, and adaptive calibration of the decision threshold.

The same pipeline, with minor adjustments to the filterbank and analysis window parameters, proves effective for detecting and identifying drones by the acoustic fingerprint of their motors. The technical convergence between the two domains opens opportunities for knowledge transfer, shared datasets and unified multi-task architectures.

As UAVs proliferate in urban and peri-urban airspace, lightweight acoustic detection systems, deployable on low-cost edge devices, will become an essential complementary element of the aerial security ecosystem – and the knowledge accumulated in speech processing represents technical capital that transfers directly to this new field of applications.

REFERENCES

[1] Park, T.J., Kanda, N., Dimitriadis, D., Han, K.J., Watanabe, S., & Narayanan, S. (2022). "A review of speaker diarization: Recent advances with deep learning." Computer Speech & Language, 72, 101317. https://doi.org/10.1016/j.csl.2021.101317

[2] Dehak, N., Kenny, P.J., Dehak, R., Dumouchel, P., & Ouellet, P. (2011). "Front-End Factor Analysis for Speaker Verification." IEEE Transactions on Audio, Speech, and Language Processing, 19(4), 788–798. https://doi.org/10.1109/TASL.2010.2064307

[3] Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018). "X-Vectors: Robust DNN Embeddings for Speaker Recognition." Proc. ICASSP, Calgary, Canada, pp. 5329–5333.

[4] Jeon, H., Kim, Y., Kim, D., Kim, J., & Kim, J. (2017). "Empirical Study of Drone Sound Detection in Real-Life Environment with Deep Neural Network." Proc. 25th European Signal Processing Conference (EUSIPCO), Kos, Greece, pp. 1858–1862.

[5] Al-Emadi, S., Al-Ali, A., Mohammad, A., & Al-Ali, A. (2019). "Audio-Based Drone Detection and Identification Using Deep Learning." Proc. 2019 15th International Wireless Communications & Mobile Computing Conference (IWCMC), pp. 459–464.

[6] Davis, S.B., & Mermelstein, P. (1980). "Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences." IEEE Transactions on Acoustics, Speech, and Signal Processing, 28(4), 357–366. https://doi.org/10.1109/TASSP.1980.1163420

[7] Reynolds, D.A., Quatieri, T.F., & Dunn, R.B. (2000). "Speaker Verification Using Adapted Gaussian Mixture Models." Digital Signal Processing, 10(1–3), 19–41. https://doi.org/10.1006/dspr.1999.0361

[8] Anguera, X., Bozonnet, S., Evans, N., Fredouille, C., Friedland, G., & Vinyals, O. (2012). "Speaker Diarization: A Review of Recent Research." IEEE Transactions on Audio, Speech, and Language Processing, 20(2), 356–370. https://doi.org/10.1109/TASL.2011.2125954

[9] Smus, B. (2013). Web Audio API: Advanced Sound for Games and Interactive Apps. O'Reilly Media. ISBN: 978-1-449-33268-5.

[10] MDN Web Docs. (2024). "Web Audio API." Mozilla Developer Network. https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API

2 Comments

0 votes
0
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Why Gaming Voice Chat Apps Are the New Way to Meet Teammates

jacksparrow - Sep 18

5 Web Dev Pitfalls That Are Silently Killing Your Projects (With Real Fixes)

Dharanidharan - Mar 3

How to save time while debugging

Codeac.io - Dec 11, 2025

Apollo Pricing: Plans, Credits, and Real Small-Team Cost

stepan-nikonov - Aug 29

Slowing AI Development Could End America’s Lead Over China: US Speaker

YasirAwan4831 - Sep 17
chevron_left
678 Points • 11 Badges
2Posts
2Comments
9Connections
I am a technology leader and entrepreneur with extensive experience in software engineering and prod... Show more

Related Jobs

View all jobs →

Commenters (This Week)

1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!