Fiverr transcription SEO for captions and transcripts

Clips live or die on captions, and every caption file starts as a transcript. Yet buyers search for three outcomes under one word: a verbatim transcript for the archive, a cleaned-up read for publication, and an SRT or VTT file ready to upload. A gig that promises transcription only leaves them guessing which one arrives.

This guide separates the three deliverables, sets the speaker-label and timestamp conventions buyers expect, shows how to prove accuracy without handing over sensitive audio, and prices the work per audio minute. The keyword rows below are examples to validate on your own account rather than platform data.

Free plan available. Local-first data. Human review on every change.

Transcription deliverable map showing verbatim transcripts, clean reads, and SRT or VTT caption files with timestamp conventions
Three formats, one honest accuracy claim

Verbatim, clean read, or caption file

Transcription covers three products that are checked, priced, and delivered differently. A verbatim transcript records everything said, including false starts and overlaps, for research, legal, and archival buyers. A clean read removes stumbles and filler while keeping the meaning. A caption file adds timing so the text can be uploaded to a player or burned in.

Fiverr's public Transcription category page, checked October 1, 2026, lists Subtitles and Captions beside language subcategories including English transcription, Arabic transcription, French transcription, and Spanish transcripts. Buyers expect the format to be named, so pick one deliverable per gig instead of promising all three.

The three transcription deliverables (illustrative framing)
DeliverableWhat lands in the folderTypical buyer
Verbatim transcriptFull text with every utterance kept and speaker turns markedResearch, legal, and archival buyers
Clean readEdited text with filler removed and meaning intactPodcasters and content teams
Caption fileSRT or VTT timed to the mediaCreators publishing short-form video
Text plus captionsTranscript and the timed file in two formatsBuyers repurposing one recording

Keyword examples to validate locally

Build the list from the audio you have delivered, then test each phrase on Fiverr: type it, read the autocomplete, open the gigs ranking for it, and note the format and turnaround they promise. You are reading how buyers phrase the job, not estimating demand.

The set below fits one seller who works in English with podcast and creator audio. Run each phrase through live search before a gig is built around it. Fiverr's Advanced Analytics article describes the order the algorithm reads, title first, then the positive keywords field, then the description, and it asks for at least 30 days before you judge new keywords.

Example keyword set for one transcriber (examples to validate locally, not platform data)
SlotExample phraseWhy a buyer types it
Primaryfiverr transcriptionThe category phrase in a buyer's own words
Secondarycaptioning services fiverrOutcome named before the platform
Secondarysubtitle fiverrShort-form creator scanning for uploads
Long-tailenglish audio transcriptionLanguage scope stated up front
Long-tailsrt caption file for videoThe exact file the buyer needs

Speaker labels, timestamps, and file formats

Conventions make a file usable, and buyers should not have to ask. Decide three things before you publish: how speakers are labeled, whether you stamp at every cue or at an interval, and which formats you deliver. Names supplied by the buyer come first; otherwise a neutral Speaker 1, Speaker 2 scheme keeps the file honest.

Format basics belong in the description so the buyer can check fit in one read. A SubRip file runs numbered cues with millisecond timing, such as 00:00:12,400 to 00:00:15,100 followed by the spoken line. A WebVTT file opens with a WEBVTT header and carries the same cue text. State which one you produce and how speakers are labeled.

  • Speaker labels — the buyer's names, or Speaker 1, Speaker 2 when none were given
  • Timing — one stated interval, written the same way in every cue
  • SRT — sequential cue numbers, millisecond ranges, plain text lines
  • VTT — WEBVTT header first, then the same cue text for the web player

Accuracy proof without sensitive audio

Client recordings rarely leave the order, so build the proof from material you own. A test clip you record yourself, a public-domain interview, or a short sample the buyer sends before ordering all work: deliver the transcript and the matching caption file for that clip and keep the corrections visible in the file.

Describe the checking pass in plain language: a second listen against the draft, a spelling pass for names and terms, and a timing read so captions do not drift. Fiverr's portfolio article allows one to twenty projects, each holding one to five files with a description of 120 to 1,400 characters: a transcript, a caption file, and a note about the hard passage.

  1. Pick a clip you may publish

    a short recording you made or a public-domain source

  2. Deliver both formats

    the transcript and the matching SRT or VTT file

  3. Show the fixes

    list what you corrected on the second listen

  4. State the method

    what you re-listen to and how you check the timing

Price per audio minute, with rules stated

The unit here is the audio minute, and the honest version says what counts: raw runtime or trimmed file, one speaker or several, whether timestamps are included, and how a rush window is priced. Fiverr's gig creation article states you can offer three packages with their own delivery times, revision counts, and prices, that at least one revision option is required, and that the minimum starting price is $5, though some categories set a higher one.

As an illustrative example: 45 minutes of single-speaker narration delivered as plain text is a different job from 45 minutes of a two-speaker interview delivered as a clean read plus an SRT, even though the runtime matches. Quote them as separate bands. The ladder below is illustrative structure only; set the minutes and windows from your own timing.

Illustrative per-audio-minute package ladder (examples, not observed rates)
TierRuntime exampleDelivery example
BasicUp to 30 minutes, single speaker, text only2 days
StandardUp to 60 minutes, clean read with speaker labels3 days
PremiumUp to 90 minutes, transcript plus SRT or VTT5 days
RushAny band, window agreed before the orderNext-day slot

Turnaround when a clip is going viral

A creator sends a 90-second clip late in the day and wants captions before the morning, because the trend will be over by lunch. That order is not a podcast backlog, and sellers who price every file at three days either refuse the fast work or accept it and miss the date.

Separate turnaround from volume: sell the fast window as its own package or extra, state the hours you need, and stop taking rush clips in hours you cannot work. Fiverr's order guide states that an order marked as Late affects your on-time delivery metric, so the window you quote is measured.

Transcription mistakes that lose repeat orders

The first mistake is no language or accent scoping. English transcription covers a fast-talking panel show and a carefully enunciated audiobook, and a gig that claims both gets orders it cannot deliver well. Name the language, the accent range, and the subject matter you can follow.

The second is an unclear timestamp format. If the description never says whether cues arrive every few seconds or only at speaker changes, buyers import a file their editor rejects and ask for a rewrite. The third is ignoring turnaround while chasing short-form work; the short form video editing guide shows the pace that lane expects.

Say the format, the language scope, and the window in the listing; buyers comparing transcripts read those lines first. The podcast editors guide covers the audio work that usually arrives alongside yours.

Where Seller OS helps

Seller OS keeps the commercial side of a transcription gig in order. The Gig Builder turns a one-line idea, such as English interviews with SRT delivery, into a structured listing with title, five tags, description, and three packages, and fills the Fiverr wizard while you review each field. Package pricing review then checks your bands against your recent applied prices and profit assumptions.

Inbox reply drafts prepare the runtime and format answer before you send it, and client records hold each buyer's naming and timing preferences for the next file. It drafts and fills; a person performs every final Send, Save, Continue, or Publish action, and it never transcribes your audio.

Seller OS package pricing review showing transcription tier bands against recent applied prices
Bands checked before the package goes live

Fiverr SEO for Transcription and Captioning questions

Should I sell verbatim or clean read transcription on Fiverr?

Sell the one you can demonstrate. Verbatim work is judged on completeness and priced for the difficulty of the speakers; a clean read is judged on readability. Put the type in the title so the right buyer finds it, describe the editing you apply, and add the second type as a package once the first has reviews.

What timestamp format do caption buyers expect?

There is no single platform rule, so state yours. Most buyers ask for SRT or VTT files, and each editor has its own cue expectations. Say which format you deliver, how often cues appear, and how speakers are labeled, so the buyer can confirm fit before ordering.

How do I price transcription per audio minute?

Set bands by runtime and job type rather than one flat rate: text only, clean read with labels, or transcript plus captions. State what counts as a minute, how multi-speaker audio is priced, and what a rush window costs. Quote from your own history so the band survives a real file.

How do I prove transcription accuracy without client files?

Record or choose a clip you may publish, deliver the transcript and the caption file for it, and show what you corrected on a second listen. A test clip, a public-domain sample, or a short file the buyer sends beforehand all answer the accuracy question without exposing anyone's audio.

Name the file, then price the minutes

Pick one deliverable, publish a test clip, and set runtime bands a buyer can quote against in a single message.