Buffered Whisper Transformer: A Real-time Speech Recognition System for Edge Devices via Dynamic Audio Chunking Mechanism
By Fei PAN, Yechuan ZHOU, Junhao ZHANG, Yuchen TAN
Highlights
Empowers non-streaming recognition models with accurate real-time speech recognition capabilities
Dynamic audio chunking with buffering queue enabling uninterrupted processing and continuous result updates
Effectively mitigates semantic fragmentation common in commercial streaming recognition systems
Robust multilingual and dialect recognition
Applications
Unlocks broad applications demanding high-precision recognition, such as simultaneous interpretation and live captioning
Enables smart home and in-vehicle voice system solutions with improved responsiveness and reliability