Newsletter
Join the Community
Subscribe to our newsletter for the latest news and updates
AI speech to text workspace that transcribes audio, video, live recordings and media URLs into editable, exportable transcripts, powered by OpenAI Whisper.
Traffic, search & AI signals for speechtext.co.
Third-party traffic estimate · Updated Sep 9, 2026
Submit your own product to reach creators and founders looking for the next tool to try.
Speech Text is an online AI speech to text workspace that turns audio, video, live browser recordings, and media URLs into editable, searchable transcripts. Powered by OpenAI Whisper, it supports more than 100 languages and runs entirely in the browser with no setup.
Speech Text is a browser-based speech to text tool that transcribes audio and video into text you can search, edit, and export. It accepts file uploads (MP3, WAV, M4A, MP4, MOV, WEBM), live browser recordings, and media URLs, and produces transcripts with timestamps and speaker labels. The service uses OpenAI Whisper technology and processes transcription in the browser via WebGPU, Transformers.js, and ONNX Runtime.
Journalists can transcribe interviews and convert them into searchable notes. Podcasters can turn episode audio into drafts for show notes, captions, or blog posts. Video editors can create SRT subtitle files from video tracks. Students and researchers can convert lectures into searchable transcripts for studying.
Start a transcription job by uploading a file, recording live, or pasting a media URL. The audio is processed in the browser using OpenAI Whisper, and the transcript appears in an editor where you can search, correct, and add speaker labels. When ready, export the transcript in TXT, SRT, VTT, DOCX, JSON, or PDF format.
Speech Text uses a freemium model. New users receive 5 free transcription minutes. Paid plans offer yearly billing with monthly equivalents:
Yes, Speech Text gives new users 5 free transcription minutes to test the workflow. After that, you can subscribe to Starter, Pro, or Max plans, or buy credit packs for extra transcription minutes.
Speech Text supports over 100 languages, including English, Spanish, Chinese, French, German, Portuguese, Russian, Korean, Japanese, and many more. It can auto-detect the language or use your manual selection.
Yes. You can upload video files in MP4, MOV, or WEBM format, or paste a media URL. Speech Text extracts the spoken track and converts it into an editable transcript, which you can export as SRT subtitles or other formats.
Speech Text exports transcripts as TXT, SRT, VTT, DOCX, JSON, and PDF. This covers plain text, subtitles, documents, and structured data for downstream applications.
Accuracy depends on source quality. Clear audio with minimal background noise, single speakers, and standard vocabulary produces the best results. The page notes that important names and facts should always be reviewed.