Speech-to-text comparisons
Fish Audio vs ElevenLabs Scribe v2
Real multi-speaker recordings, transcribed by both. Play the audio and follow the two transcripts side by side, with speakers and word timings.
Internal sales demo for the Fish Audio team. Please do not share these pages or links outside Fish Audio.
-
EN
English meeting
English · 4 speakers · 15:00
Four people in a casual chat, with crosstalk and laughter.
Open the comparison → -
JA
Japanese conversation
Japanese · 2 speakers · 9:32
Two people chat about YouTubers, video games and car costs.
Open the comparison → -
ES
Spanish conversation
Spanish · 2 speakers · 14:43
Two people talk about drawing and digital art.
Open the comparison → -
EN
English customer-service call
English · 2 speakers · 5:59
A caller asks a hotel reservations agent for room rates, on phone-line audio.
Open the comparison →
About these pages
Recordings for the meeting, Japanese and Spanish pages: Basis Conversations 1500, published by Basis (https://huggingface.co/datasets/basis-ai/basis-conversations-1500, revision 05a19e06). The customer-service call is a real call from internal contact-centre data.
Method: each recording was transcribed once by each service from the same audio file; no language or number of speakers was given. Fish Audio: model transcribe-1-pro (the production build, on the production model servers), default settings apart from word timestamps; it is not the default model of /v1/asr, so select it by name to reproduce these results. ElevenLabs: Scribe v2, default settings apart from speaker diarization, audio-event tags and word timestamps. Transcripts are shown as returned by each service, and each page names the date of its run. One recording is not a benchmark.