Apple's New SpeechAnalyzer API Crushes OpenAI's Whisper Small in English Benchmark, Delivers 3x Faster Processing

Apple's on-device speech recognition API "SpeechAnalyzer," designed for iOS 26 and macOS Tahoe 26, has outperformed OpenAI's Whisper Small in its first third-party benchmark. The development team behind the transcription app Inscribe tested the API using the English dataset LibriSpeech, finding that SpeechAnalyzer achieved a word error rate of 2.12% on clear speech, beating Whisper Small's 3.74%, while delivering approximately three times the processing speed. Compared to the legacy SFSpeechRecognizer, the error rate was reduced by roughly 75%. Following these results, Inscribe modified its engine selection feature to prioritize SpeechAnalyzer for supported languages. However, the validation was limited to English read-aloud speech, leaving performance in real-world scenarios such as meetings or distant recordings as a subject for future evaluation.
Apple's New SpeechAnalyzer API Crushes OpenAI's Whisper Small in English Benchmark, Delivers 3x Faster Processing

Apple's on-device speech recognition API "SpeechAnalyzer," set to debut in its next-generation operating systems, has had its capabilities revealed through its first third-party benchmark. Testing conducted by the development team behind the transcription app "Inscribe" found that SpeechAnalyzer surpassed the accuracy of OpenAI's "Whisper" Small model while achieving approximately three times the processing speed.

This benchmark marks the first quantification of the specific performance of SpeechAnalyzer, which Apple announced at its 2025 Worldwide Developers Conference (WWDC25). Apple is providing the API for iOS 26 and macOS Tahoe 26, touting its ability to handle long recordings, conversations, and audio captured from a distance. However, the company had not previously disclosed specific accuracy figures compared to the legacy SFSpeechRecognizer or competing models.

To establish selection criteria for the speech recognition engine integrated into their app, Inscribe's development team compared five engines under identical conditions using "LibriSpeech," a standard dataset for English speech recognition. The evaluation metric was Word Error Rate (WER), where a lower value indicates higher recognition accuracy.

The test results showed that for 2,620 samples of clear read-aloud speech, SpeechAnalyzer achieved the lowest WER at 2.12%, significantly outperforming Whisper Small at 3.74%, Whisper Base at 5.42%, and Whisper Tiny at 7.88%. The legacy SFSpeechRecognizer recorded a 9.02% error rate, more than four times that of SpeechAnalyzer.

The gap widened further in 2,939 samples of noisy, challenging read-aloud speech. SpeechAnalyzer's WER was held to 4.56%, while Whisper Small reached 7.95%, and SFSpeechRecognizer hit 16.25%—roughly 3.6 times higher than SpeechAnalyzer. Inscribe's development team analyzed that compared to SFSpeechRecognizer, SpeechAnalyzer reduced the word error rate by approximately 75%.

SpeechAnalyzer held an advantage not only in accuracy but also in processing speed. All tests were executed on a Mac equipped with an M2 Pro chip and 32GB of memory, without sending audio data to external servers. The audio processing time per second for SpeechAnalyzer was reportedly about one-third that of Whisper Small. While a one-hour audio file could be processed by any engine in roughly 1.5 to 5 minutes, it should be noted that development processes were running concurrently on the same Mac during the benchmark, introducing variability in processing times. Inscribe has indicated plans to conduct a re-measurement with the Mac in an idle state and update the results.

Following these results, Inscribe modified the automatic speech recognition engine selection feature in its app. The design was switched to prioritize the SpeechAnalyzer API for supported languages, while falling back to Whisper for unsupported languages.

However, this benchmark has important limitations. The validation was restricted to English read-aloud speech, and it remains unconfirmed whether similar results can be achieved in practical scenarios such as accented speakers, meetings with multiple simultaneous speakers, or distant recordings. At WWDC25, where Apple unveiled SpeechAnalyzer, the company explained that the same recognition technology is integrated into apps like Notes, Voice Memos, and Journal, but performance in diverse usage environments awaits further verification.

Meanwhile, Apple's software strategy extends beyond speech recognition. The company has begun rolling out public beta versions of its 2027 OS updates: iOS 27, iPadOS 27, macOS 27 Golden Gate, and watchOS 27. The biggest highlight is "Siri AI," which has evolved from a basic voice command system into a modern AI assistant. It enables natural conversation, on-screen content recognition, and multi-step actions within apps, while Apple Intelligence-compatible devices also gain the ability to customize vocal expressiveness.

Siri AI is currently available only in English and is initially restricted within the European Union (EU), though a dedicated app will save interaction history. Beyond AI features, iOS 27 also delivers broad performance improvements, including up to 30% faster app launch speeds, up to 70% faster photo display, and up to 80% faster AirDrop transfer speeds.

The high baseline performance of SpeechAnalyzer could become a critical component supporting Apple's multimodal AI strategy. Its high accuracy and speed, specialized for on-device processing, position it as a foundational technology enabling Siri AI to understand speech in real-time while achieving end-to-end privacy protection on the device.

Add to Google Preferred Sources

Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.







More Related News