Skip to content
IronMemo
Multilingual transcription

Transcription in 30+ languages, not one model for all of them

IronMemo transcribes meetings and recordings in 30+ languages. Instead of running one universal engine on everything, we scored 95 speech models against 120+ independent sources and route each language to the model that performed best for it, with a second model on standby. English is the easy part - the difference shows up everywhere else.

Updated July 25, 2026

  • 30+ languages
  • 95 models benchmarked
  • Files up to 2 GB

Languages we transcribe

Each card names what makes that language hard for speech recognition - which is precisely why one universal model cannot be the best at all of them.

en

English

English

The best-covered language in every engine. What still breaks it: crosstalk, strong regional accents and domain jargon.

es

Spanish

Español

Fast delivery and wide regional variation - Rioplatense, Caribbean and Castilian differ in rhythm, vocabulary and pronoun use.

de

German

Deutsch

Compound nouns written without spaces, and verbs parked at the end of long subordinate clauses.

fr

French

Français

Liaison and elision erase word boundaries - connected speech runs several words together into one sound.

pt

Portuguese

Português

Brazilian and European variants diverge sharply, and European Portuguese swallows unstressed vowels almost entirely.

ja

Japanese

日本語

No spaces between words, three scripts mixed in one sentence, and pitch accent that changes meaning.

zh

Chinese (Mandarin)

中文

Tone carries meaning, and written Chinese offers no word boundaries to fall back on.

ko

Korean

한국어

Agglutinative endings plus speech levels - the same verb changes shape depending on who is speaking to whom.

it

Italian

Italiano

Vowel-final words and fast, overlapping delivery blur where one word ends and the next begins.

nl

Dutch

Nederlands

Guttural consonants, and business meetings that switch into English mid-sentence out of pure habit.

pl

Polish

Polski

Dense consonant clusters and seven cases, with palatalization shifting sounds inside the word.

tr

Turkish

Türkçe

Agglutination: a single Turkish word can carry what English needs an entire clause to say.

hi

Hindi

हिन्दी

Constant English code-switching - Hinglish is ordinary business speech here, not an edge case.

ar

Arabic

العربية

Modern Standard Arabic and the spoken dialects differ more than some European languages do, and short vowels are not written at all.

IronMemo transcribes 30+ languages. The fifteen named above - English, Spanish, Russian, German, French, Portuguese, Japanese, Chinese (Mandarin), Korean, Italian, Dutch, Polish, Turkish, Hindi and Arabic - cover the bulk of business meetings, and the remaining languages run through the same pipeline. Updated July 2026.

When one meeting has two languages

Most notetakers make you pick a language up front. Here is what that means in practice, in their own words.

What most tools do

  • Read.ai states plainly that meetings where several languages are spoken are not fully supported and may produce inconsistent results.
  • Circleback detects the predominantly spoken language and transcribes the entire meeting in it - the second language is processed as if it were the first.
  • Krisp asks you to select one language up front, and the transcript is generated in that language whatever is actually said.
  • Fireflies does detect switches at word level, but its multi-language mode is limited to the Business and Enterprise plans.

What IronMemo does

  • Language is detected as the conversation goes, so a switch partway through does not restart or corrupt the transcript.
  • Russian-English and Spanish-English switching inside a single sentence is a case we tuned for, not one we merely tolerate.
  • Speaker labels survive the switch - diarization does not reset when the language changes.
  • Every line keeps its timestamp, so you can click back into the audio and check any sentence yourself.

Fair warning: mid-sentence switching is genuinely hard for every engine, and Sonix says so publicly - their detection works best when languages change between speakers rather than mid-sentence. If your meetings run in one widely supported language, most tools listed here will serve you well; the difference appears when they do not. Vendor statements checked in July 2026.

How a language gets its model

The short version. The full methodology, including which models we banned and why, lives on our AI page.

  1. Benchmark first, marketing never

    95 speech models scored against 120+ independent sources - accuracy, behavior on silence, diarization quality and cost, measured per language rather than in aggregate.

  2. Route each language to its winner

    Fourteen languages are pinned to a primary model with a second on standby. The rest run on the strongest general-purpose model from the same benchmark.

  3. Keep the result auditable

    Every transcript line stays clickable back to its own second of audio, so an accuracy claim is something you verify yourself rather than take on trust.

Counting languages is the easy part

Language totals exactly as each vendor advertises them - next to what actually happens to an individual language.

ServiceLanguages advertised
IronMemo30+
HappyScribe150+
Fireflies100+
Notta58
Sonix54+

Totals are what each vendor advertised on its own pages in July 2026: happyscribe.com/languages, the Fireflies help center, notta.ai pricing and sonix.ai/languages. A larger number is not a worse product - if your meetings run in one widely supported language, several of these will do the job. The count stops being the answer when your language is the one handed to a generic model.

What the benchmark actually is

The numbers behind per-language routing - a research database, not a slogan.

95

speech models evaluated

Scored on accuracy, behavior on silence, diarization and cost before any of them ever touched a customer recording.

120+

independent sources

Academic benchmarks, vendor documentation and our own test corpus, with an evidence grade attached to every result.

14

languages benchmarked in depth

Each with a primary model and a fallback. The remaining languages run on the strongest general model in the same benchmark.

Questions about language support

30+ languages. The fifteen named on this page - English, Spanish, Russian, German, French, Portuguese, Japanese, Chinese (Mandarin), Korean, Italian, Dutch, Polish, Turkish, Hindi and Arabic - cover the bulk of business meetings, and the rest run through the same pipeline. If your language is not on the list, ask before you sign up and we will tell you honestly how it scores.

No, and it would be dishonest to claim otherwise. Fourteen languages are benchmarked in depth: candidate models are tested on that specific language and a primary plus a fallback is pinned. Every other language is routed to the strongest general-purpose model in the same benchmark. So each language gets the best engine we measured for it - what differs is how deep the measurement went. Deep-benchmarked languages are the ones that get a dedicated page, starting with Russian.

The transcript follows the switch instead of forcing the meeting into a single language, and speaker labels stay intact. That is worth checking against alternatives: Read.ai documents that multi-language meetings are not fully supported, Circleback transcribes the whole meeting in whichever language dominates, and Fireflies restricts word-level switch detection to its Business and Enterprise plans. Checked against vendor documentation in July 2026.

It is different and it is measurable. Notta publishes a single up-to-98.86% figure, Sonix quotes 85-99% depending on audio quality, and Transkriptor claims 99% across all 100+ of its languages - one number for every language each of them supports. No speech model performs identically on English and on Turkish or Arabic. We publish word error rates language by language on our accuracy page instead of one headline figure.

Not today. Routing is automatic and driven by the benchmark rather than by guesswork, and the fallback engages when the primary model struggles with a given recording. What you can inspect is the result: every line is clickable back to the exact second of audio it came from, so you can verify any sentence rather than trust an accuracy percentage.

Ready to stop losing what matters in your meetings?

Start for free. No credit card required.

Free plan · No credit card · Setup in 60 seconds