Speech-to-text comparisons
Internal

Speech-to-text comparisons

Fish Audio vs ElevenLabs Scribe v2

Real multi-speaker recordings, transcribed by both. Play the audio and follow the two transcripts side by side, with speakers and word timings.

Internal sales demo for the Fish Audio team. Please do not share these pages or links outside Fish Audio.

About these pages

Recordings for the meeting, Japanese and Spanish pages: Basis Conversations 1500, published by Basis (https://huggingface.co/datasets/basis-ai/basis-conversations-1500, revision 05a19e06). The customer-service call is a real call from internal contact-centre data.

Method: each recording was transcribed once by each service from the same audio file; no language or number of speakers was given. Fish Audio: model transcribe-1-pro (the production build, on the production model servers), default settings apart from word timestamps; it is not the default model of /v1/asr, so select it by name to reproduce these results. ElevenLabs: Scribe v2, default settings apart from speaker diarization, audio-event tags and word timestamps. Transcripts are shown as returned by each service, and each page names the date of its run. One recording is not a benchmark.