Open benchmark · measured 2026-08-24

Spanish over the phone

Most speech-to-text rankings measure studio-read English. Your audio is a two-speaker phone call in Latin American Spanish. We measured seven engines on 13 real calls, linked every audio file, and published the result — including the part where we come fourth.

13
real phone calls, five dialects
7
engines measured
100%
of the audio is public
47 pts
minimum spread per engine, same language

The results

WER is the share of words the engine got wrong — lower is better. Attribution is the share of words assigned to the correct speaker — higher is better. They are not interchangeable, and the ranking flips depending on which one you care about.

engineWERattributionbest on
Soniox20.01%94.4%6 / 13
AssemblyAI20.45%93.9%4 / 13
Speechmatics21.48%86.8%0 / 13
Orchard23.15%89.5%3 / 13
Gladia25.13%90.9%0 / 13
Deepgram26.81%92.6%0 / 13
Rev.ai30.45%89.6%0 / 13

ElevenLabs is excluded: it completed 9 of the 13 calls before hitting a quota limit. On those nine it scored 24.9% WER and 95.1% attribution.

Call by call

Every row links to the original audio on YouTube. Green marks the best engine on that call; no engine wins everywhere — the leader takes 6 of 13.

callcountryOrchardSonioxAssemblyAISpeechmaticsGladiaDeepgramRev.ai
ar_RfaVDSUKjjQArgentina22.6%11.0%16.3%12.0%20.9%17.3%21.9%
ar_c-2azay9ppUArgentina51.1%38.6%40.1%41.4%73.7%61.1%69.8%
co_OMBeyGNCWDYColombia11.4%6.3%6.5%8.5%8.5%11.9%9.1%
co_arP03vXvhmcColombia60.9%54.6%56.7%56.0%57.8%63.2%61.7%
co_wqICQbLdgTQColombia44.2%42.5%41.6%43.1%44.6%47.0%49.1%
mx_a6tvavirPVgMexico24.0%27.9%26.7%26.4%25.5%38.2%36.5%
mx_hdwL9eiNkT8Mexico14.1%9.9%12.9%14.2%15.9%14.4%22.1%
pe_h2utUsdQi5kPeru8.3%6.9%6.6%10.9%8.0%11.8%19.4%
call_center_ytCall center13.4%11.2%13.4%13.5%14.1%15.7%18.3%
call_Evf5ZjqEE4oArgentina12.8%12.0%10.9%15.4%12.8%17.6%27.4%
call_9eo8SdXdArQColombia14.7%12.3%9.2%10.4%12.9%20.9%22.7%
call_G0KisoxpgnAColombia7.5%7.7%8.4%10.0%15.5%11.5%12.9%
call_7X01soGw9YAColombia15.9%19.4%16.5%17.4%16.5%17.9%24.9%

Method, in five sentences

  • References were hand-transcribed or taken from each video's own subtitles; the hypothesis is clipped to the union of reference turns (±0.5 s), identically for every engine.
  • Text is normalized before comparison: lowercase, no accents or punctuation, loose digits joined.
  • Nine references are published in the repo. Four are a held-out test set: we publish the audio, the turn counts and a hash of each text — not the text — so no engine can train on them. Send us your transcript of those four calls and we score it with the same code, whoever wins.
  • Two audios were excluded because their reference covers under 20 words per minute of continuous speech, plus one TV scene that is not a call.
  • Every number on this page regenerates from one CSV in the repo. If you find an error, open an issue — we would rather fix it than keep a pretty number.

Audio links, public references, scoring code and raw results:

github.com/Orchard-Run/orchard-stt-bench

Fourth place, three points off the leader, in the engine's first week — measured against companies that raised between $86M and $158M. The short-term goal is first place in Spanish, and this page will keep updating either way. Try the API.