Transcribe hours of audio with Whisper
Whisper large-v3 turns recordings into text in about 100 languages, including Nepali and Hindi. On a GPU it runs many times faster than real time, so a hundred hours of audio becomes an afternoon’s work instead of a week’s.
Upload your files, run Whisper over the whole folder, and download plain text or subtitle files.
Which GPU to pick
Whisper large-v3 fits comfortably on a 12–16 GB card. A 24 GB card lets you process more audio at once for higher throughput.
Transcribe 100 hours of audio (Whisper large-v3, batched) costs, on today’s cheapest hosts:
| GPU | Typical time | Cost today |
|---|---|---|
| RTX 3090 | 4.5 h | USD 1.04 |
| A100 SXM4 | 2 h | USD 1.56 |
| RTX 4090 | 3 h | USD 1.74 |
| RTX A5000 | 5 h | USD 1.90 |
| L40S | 2.5 h | USD 2.15 |
Ready-made software
Pick one of these when you rent and it starts with everything installed:
- Whisper speech-to-text — PyTorch image with faster-whisper and ffmpeg installed on first boot. For Nepali/English transcription jobs.
How it works
- Create an account with Google or your email address. There is no ID upload: your first payment is how we know who you are.
- Add credits in your own currency.
- Add your SSH key, then pick the software and a GPU from the live board.
- Wait for Running — about a minute on a warm host, a few minutes on a cold one — then open the app or connect over SSH.
- Stop when you are done. GPU billing ends at once; only the disk is billed until you destroy the instance.
Questions
How accurate is Whisper on Nepali?
Usable for drafts and search, weaker than on English, and best on clear recordings. Phone-call audio and heavy background noise lower accuracy noticeably.
Can I process a whole folder automatically?
Yes. Run Whisper from a notebook or the command line over every file in the folder and collect the transcripts.
Is my audio kept anywhere?
We do not access your instance or your files. When you destroy the instance, its disk is deleted.