Open benchmark · measured 2026-08-24
Spanish over the phone
Most speech-to-text rankings measure studio-read English. Your audio is a two-speaker phone call in Latin American Spanish. We measured seven engines on 13 real calls, linked every audio file, and published the result — including the part where we come fourth.
The results
WER is the share of words the engine got wrong — lower is better. Attribution is the share of words assigned to the correct speaker — higher is better. They are not interchangeable, and the ranking flips depending on which one you care about.
| engine | WER | attribution | best on |
|---|---|---|---|
| Soniox | 20.01% | 94.4% | 6 / 13 |
| AssemblyAI | 20.45% | 93.9% | 4 / 13 |
| Speechmatics | 21.48% | 86.8% | 0 / 13 |
| Orchard | 23.15% | 89.5% | 3 / 13 |
| Gladia | 25.13% | 90.9% | 0 / 13 |
| Deepgram | 26.81% | 92.6% | 0 / 13 |
| Rev.ai | 30.45% | 89.6% | 0 / 13 |
ElevenLabs is excluded: it completed 9 of the 13 calls before hitting a quota limit. On those nine it scored 24.9% WER and 95.1% attribution.
Call by call
Every row links to the original audio on YouTube. Green marks the best engine on that call; no engine wins everywhere — the leader takes 6 of 13.
| call | country | Orchard | Soniox | AssemblyAI | Speechmatics | Gladia | Deepgram | Rev.ai |
|---|---|---|---|---|---|---|---|---|
| ar_RfaVDSUKjjQ | Argentina | 22.6% | 11.0% | 16.3% | 12.0% | 20.9% | 17.3% | 21.9% |
| ar_c-2azay9ppU | Argentina | 51.1% | 38.6% | 40.1% | 41.4% | 73.7% | 61.1% | 69.8% |
| co_OMBeyGNCWDY | Colombia | 11.4% | 6.3% | 6.5% | 8.5% | 8.5% | 11.9% | 9.1% |
| co_arP03vXvhmc | Colombia | 60.9% | 54.6% | 56.7% | 56.0% | 57.8% | 63.2% | 61.7% |
| co_wqICQbLdgTQ | Colombia | 44.2% | 42.5% | 41.6% | 43.1% | 44.6% | 47.0% | 49.1% |
| mx_a6tvavirPVg | Mexico | 24.0% | 27.9% | 26.7% | 26.4% | 25.5% | 38.2% | 36.5% |
| mx_hdwL9eiNkT8 | Mexico | 14.1% | 9.9% | 12.9% | 14.2% | 15.9% | 14.4% | 22.1% |
| pe_h2utUsdQi5k | Peru | 8.3% | 6.9% | 6.6% | 10.9% | 8.0% | 11.8% | 19.4% |
| call_center_yt | Call center | 13.4% | 11.2% | 13.4% | 13.5% | 14.1% | 15.7% | 18.3% |
| call_Evf5ZjqEE4o | Argentina | 12.8% | 12.0% | 10.9% | 15.4% | 12.8% | 17.6% | 27.4% |
| call_9eo8SdXdArQ | Colombia | 14.7% | 12.3% | 9.2% | 10.4% | 12.9% | 20.9% | 22.7% |
| call_G0KisoxpgnA | Colombia | 7.5% | 7.7% | 8.4% | 10.0% | 15.5% | 11.5% | 12.9% |
| call_7X01soGw9YA | Colombia | 15.9% | 19.4% | 16.5% | 17.4% | 16.5% | 17.9% | 24.9% |
Method, in five sentences
- References were hand-transcribed or taken from each video's own subtitles; the hypothesis is clipped to the union of reference turns (±0.5 s), identically for every engine.
- Text is normalized before comparison: lowercase, no accents or punctuation, loose digits joined.
- Nine references are published in the repo. Four are a held-out test set: we publish the audio, the turn counts and a hash of each text — not the text — so no engine can train on them. Send us your transcript of those four calls and we score it with the same code, whoever wins.
- Two audios were excluded because their reference covers under 20 words per minute of continuous speech, plus one TV scene that is not a call.
- Every number on this page regenerates from one CSV in the repo. If you find an error, open an issue — we would rather fix it than keep a pretty number.
Audio links, public references, scoring code and raw results:
github.com/Orchard-Run/orchard-stt-bench
Fourth place, three points off the leader, in the engine's first week — measured against companies that raised between $86M and $158M. The short-term goal is first place in Spanish, and this page will keep updating either way. Try the API.