On-Device Multi-Type Disfluency Detection with Sub-Millisecond Inference on Apple Silicon

Nazar Kozak

Engineering Archive preprint, April 2026 · DOI: 10.31224/6814

Abstract

This deployment-oriented preprint evaluates a multi-type speech-disfluency classifier running on Apple Silicon. It reports episode-grouped evaluation on SEP-28K, Core ML model sizes and accuracy, inference measurements across four Apple hardware generations, and numerical verification of the PyTorch-to-Core ML export path. Voice-stress features are analyzed separately as an auxiliary empirical result.

Keywords
Disfluency detection, stuttering, on-device inference, CoreML, voice stress analysis, SEP-28K, mobile speech processing
Status
Engineering Archive preprint, posted April 14, 2026
DOI
10.31224/6814
Models
DisfluencyCNN (617K parameters, 1.2 MB) · ResNet-18 Adapted (11.2M parameters, 21 MB)
Dataset
SEP-28K (20,131 clips, 5-fold episode-grouped cross-validation)
Author
Nazar Kozak — Kozak Technologies Inc., Los Angeles, CA, USA
Contact
nzrkzk@gmail.com · ORCID

← Back to all publications