Benchmark6 min read

Not the most accurate. The best accuracy per dollar for Latin American Spanish.

Fourth of seven on real Latin American calls, three points off the leader, at a fraction of the price. The math, the features that ship today, and how we size concurrency per customer instead of per plan.

Mateo Bustamante

We are not the most accurate speech-to-text engine in the world. I built it, so I get to say it. What we are is close enough to the most accurate that the difference rarely changes a business decision, at a price that changes whether the project exists at all.

That sentence is the whole pitch. The rest of this post is the evidence, including the part where we lose.

The benchmark, with the calls we lose

We measured ourselves in public against seven engines on thirteen real phone calls from Argentina, Colombia, Mexico and Peru. Real audio: sales calls, complaints, people talking over each other, the occasional terrible microphone. Every file is linked next to its number on orchardrun.com/benchmark.

We came fourth of seven, three points off the leader. The three engines above us raised between $86 million and $158 million. The leader wins six of the thirteen calls outright. We win three. Nobody wins everywhere: the ranking flips call by call, which is the first thing worth knowing before you trust any single number, including ours.

The other column

We charge $0.00042 per minute, which is $0.025 per hour of audio. With word-level speaker diarization included, $0.036 per hour. The three engines above us on the benchmark charge between four and eight times that, as published on their own pricing pages.

Three points of error rarely change a business conclusion. Paying eight times more changes whether the project exists.

That is why we say, without hedging, that we are the best accuracy per dollar for Latin American Spanish. Not the best accuracy. The best ratio. For conversation analytics, quality assurance, and anything that processes thousands of hours a month, the ratio is the number that decides the budget.

What ships today

Cheap is not the same as bare. Everything below is in production right now; nothing on this list is a roadmap item.

  • Speaker diarization at the word level, with hints for the expected number of speakers.
  • Word timestamps, for alignment and highlighting.
  • Vocabulary context: pass product names and domain jargon so the engine recognizes them.
  • Async batch with webhooks, output as JSON, SRT or VTT.
  • OpenAI-compatible API. If you integrated the Whisper API, switching is a URL change.
  • MCP server, TypeScript SDK, and an n8n kit for the no-code side.
  • On-premise under an annual license, for compliance cases.
  • We never train on customer audio. Not opt-out. Never.

Concurrency is not a plan tier

This is the part almost nobody does, and the reason most of our enterprise conversations start. Concurrency on Orchard does not come in a generic plan. We size it per customer, to the shape of their traffic.

One customer needs to process five thousand hours a month, overnight, with no hurry. Another needs ten conversations in parallel with a fifteen-second turnaround on each. Those are two different infrastructures. We build them differently, at the same price per hour. The business logic is the customer's. We adapt to it, not the other way around.

In practice that means a short conversation about volume, peak concurrency, and acceptable latency, and then a configuration that fits it. Not a slider on a pricing page.

How to check any of this

Do not take our benchmark's word for it, and do not take a competitor's either. Send us the audio that breaks transcribers today: thick accents, bad microphones, people talking over each other. Run it in parallel with whatever you use now and compare with your own numbers. There are 500 free minutes at signup, no card, and the pricing page has every competitor rate with its source linked.

If the accuracy gap costs you more than the price gap saves you, you should stay where you are. For most Spanish conversation workloads we have measured, it does not.

Try it on your own audio

500 free minutes at signup. No card, no call.