Fiverr transcription SEO for captions and transcripts
Clips live or die on captions, and every caption file starts as a transcript. Yet buyers search for three outcomes under one word: a verbatim transcript for the archive, a cleaned-up read for publication, and an SRT or VTT file ready to upload. A gig that promises transcription only leaves them guessing which one arrives.
This guide separates the three deliverables, sets the speaker-label and timestamp conventions buyers expect, shows how to prove accuracy without handing over sensitive audio, and prices the work per audio minute. The keyword rows below are examples to validate on your own account rather than platform data.
Free plan available. Local-first data. Human review on every change.

Verbatim, clean read, or caption file
Transcription covers three products that are checked, priced, and delivered differently. A verbatim transcript records everything said, including false starts and overlaps, for research, legal, and archival buyers. A clean read removes stumbles and filler while keeping the meaning. A caption file adds timing so the text can be uploaded to a player or burned in.
Fiverr's public Transcription category page, checked October 1, 2026, lists Subtitles and Captions beside language subcategories including English transcription, Arabic transcription, French transcription, and Spanish transcripts. Buyers expect the format to be named, so pick one deliverable per gig instead of promising all three.
| Deliverable | What lands in the folder | Typical buyer |
|---|---|---|
| Verbatim transcript | Full text with every utterance kept and speaker turns marked | Research, legal, and archival buyers |
| Clean read | Edited text with filler removed and meaning intact | Podcasters and content teams |
| Caption file | SRT or VTT timed to the media | Creators publishing short-form video |
| Text plus captions | Transcript and the timed file in two formats | Buyers repurposing one recording |
Keyword examples to validate locally
Build the list from the audio you have delivered, then test each phrase on Fiverr: type it, read the autocomplete, open the gigs ranking for it, and note the format and turnaround they promise. You are reading how buyers phrase the job, not estimating demand.
The set below fits one seller who works in English with podcast and creator audio. Run each phrase through live search before a gig is built around it. Fiverr's Advanced Analytics article describes the order the algorithm reads, title first, then the positive keywords field, then the description, and it asks for at least 30 days before you judge new keywords.
| Slot | Example phrase | Why a buyer types it |
|---|---|---|
| Primary | fiverr transcription | The category phrase in a buyer's own words |
| Secondary | captioning services fiverr | Outcome named before the platform |
| Secondary | subtitle fiverr | Short-form creator scanning for uploads |
| Long-tail | english audio transcription | Language scope stated up front |
| Long-tail | srt caption file for video | The exact file the buyer needs |
Speaker labels, timestamps, and file formats
Conventions make a file usable, and buyers should not have to ask. Decide three things before you publish: how speakers are labeled, whether you stamp at every cue or at an interval, and which formats you deliver. Names supplied by the buyer come first; otherwise a neutral Speaker 1, Speaker 2 scheme keeps the file honest.
Format basics belong in the description so the buyer can check fit in one read. A SubRip file runs numbered cues with millisecond timing, such as 00:00:12,400 to 00:00:15,100 followed by the spoken line. A WebVTT file opens with a WEBVTT header and carries the same cue text. State which one you produce and how speakers are labeled.
- Speaker labels — the buyer's names, or Speaker 1, Speaker 2 when none were given
- Timing — one stated interval, written the same way in every cue
- SRT — sequential cue numbers, millisecond ranges, plain text lines
- VTT — WEBVTT header first, then the same cue text for the web player
Accuracy proof without sensitive audio
Client recordings rarely leave the order, so build the proof from material you own. A test clip you record yourself, a public-domain interview, or a short sample the buyer sends before ordering all work: deliver the transcript and the matching caption file for that clip and keep the corrections visible in the file.
Describe the checking pass in plain language: a second listen against the draft, a spelling pass for names and terms, and a timing read so captions do not drift. Fiverr's portfolio article allows one to twenty projects, each holding one to five files with a description of 120 to 1,400 characters: a transcript, a caption file, and a note about the hard passage.
Pick a clip you may publish
a short recording you made or a public-domain source
Deliver both formats
the transcript and the matching SRT or VTT file
Show the fixes
list what you corrected on the second listen
State the method
what you re-listen to and how you check the timing
Price per audio minute, with rules stated
The unit here is the audio minute, and the honest version says what counts: raw runtime or trimmed file, one speaker or several, whether timestamps are included, and how a rush window is priced. Fiverr's gig creation article states you can offer three packages with their own delivery times, revision counts, and prices, that at least one revision option is required, and that the minimum starting price is $5, though some categories set a higher one.
As an illustrative example: 45 minutes of single-speaker narration delivered as plain text is a different job from 45 minutes of a two-speaker interview delivered as a clean read plus an SRT, even though the runtime matches. Quote them as separate bands. The ladder below is illustrative structure only; set the minutes and windows from your own timing.
| Tier | Runtime example | Delivery example |
|---|---|---|
| Basic | Up to 30 minutes, single speaker, text only | 2 days |
| Standard | Up to 60 minutes, clean read with speaker labels | 3 days |
| Premium | Up to 90 minutes, transcript plus SRT or VTT | 5 days |
| Rush | Any band, window agreed before the order | Next-day slot |
Turnaround when a clip is going viral
A creator sends a 90-second clip late in the day and wants captions before the morning, because the trend will be over by lunch. That order is not a podcast backlog, and sellers who price every file at three days either refuse the fast work or accept it and miss the date.
Separate turnaround from volume: sell the fast window as its own package or extra, state the hours you need, and stop taking rush clips in hours you cannot work. Fiverr's order guide states that an order marked as Late affects your on-time delivery metric, so the window you quote is measured.
Transcription mistakes that lose repeat orders
The first mistake is no language or accent scoping. English transcription covers a fast-talking panel show and a carefully enunciated audiobook, and a gig that claims both gets orders it cannot deliver well. Name the language, the accent range, and the subject matter you can follow.
The second is an unclear timestamp format. If the description never says whether cues arrive every few seconds or only at speaker changes, buyers import a file their editor rejects and ask for a rewrite. The third is ignoring turnaround while chasing short-form work; the short form video editing guide shows the pace that lane expects.
Say the format, the language scope, and the window in the listing; buyers comparing transcripts read those lines first. The podcast editors guide covers the audio work that usually arrives alongside yours.
Where Seller OS helps
Seller OS keeps the commercial side of a transcription gig in order. The Gig Builder turns a one-line idea, such as English interviews with SRT delivery, into a structured listing with title, five tags, description, and three packages, and fills the Fiverr wizard while you review each field. Package pricing review then checks your bands against your recent applied prices and profit assumptions.
Inbox reply drafts prepare the runtime and format answer before you send it, and client records hold each buyer's naming and timing preferences for the next file. It drafts and fills; a person performs every final Send, Save, Continue, or Publish action, and it never transcribes your audio.

Fiverr SEO for Transcription and Captioning questions
Should I sell verbatim or clean read transcription on Fiverr?
Sell the one you can demonstrate. Verbatim work is judged on completeness and priced for the difficulty of the speakers; a clean read is judged on readability. Put the type in the title so the right buyer finds it, describe the editing you apply, and add the second type as a package once the first has reviews.
What timestamp format do caption buyers expect?
There is no single platform rule, so state yours. Most buyers ask for SRT or VTT files, and each editor has its own cue expectations. Say which format you deliver, how often cues appear, and how speakers are labeled, so the buyer can confirm fit before ordering.
How do I price transcription per audio minute?
Set bands by runtime and job type rather than one flat rate: text only, clean read with labels, or transcript plus captions. State what counts as a minute, how multi-speaker audio is priced, and what a rush window costs. Quote from your own history so the band survives a real file.
How do I prove transcription accuracy without client files?
Record or choose a clip you may publish, deliver the transcript and the caption file for it, and show what you corrected on a second listen. A test clip, a public-domain sample, or a short file the buyer sends beforehand all answer the accuracy question without exposing anyone's audio.
Name the file, then price the minutes
Pick one deliverable, publish a test clip, and set runtime bands a buyer can quote against in a single message.