YouTube Transcript Scraper: captions, not speech-to-text
YouTube Transcript Scraper returns the caption track of a public YouTube video as structured lines, each with its text, its start time and its duration, reading the captions the video already has rather than transcribing its audio.
Free Apify plan: 5 runs a month, 10 items a run. No credit card.
- $0.001
- per transcript extracted
- $0
- for a video with no captions
- $1
- per 1,000 transcripts
- 99.95%
- run success, last 30 days
Read on 27 September 2026.
Four inputs, and a list of links it will refuse
The input surface is small and the listing is precise about its edges, which link forms work, which are rejected, and what happens to the ones in between.
startUrls$0.001per transcript
Only if it succeeds
One video link per entry, as a plain string or as { "url": "…" }. Each is processed separately, so a duplicate link returns a duplicate row.
https://www.youtube.com/watch?v=_AbFXuGDRTsA miss does not fail. It returns another track
langoptionalno charge
Check what came back
An optional caption language code, en, es, de, tr, pt-BR. Leave it empty for the video’s default track.
lang: "fr"maxItems÷ $0.001caps the bill
Run-wide
Stops the run saving once this many rows exist. Leave it empty to process every link you supplied.
100Unsupported linksrejectednot charged
watch and youtu.be only
Shorts, live, embed, playlist and channel URLs are rejected outright, as are links from other sites.
/shorts/… → rejected
Starting from a playlist or a channel? The listing points you at the Playlist and Channel Videos scrapers to collect the watch links first, then bring them here, which is exactly the pairing those two pages describe from the other end.
| Apify actor | apidojo/youtube-transcript-scraper |
|---|---|
| Inputs | startUrls · lang · maxItems · customMapFunction |
| Accepts | youtube.com/watch and youtu.be links, including m. and music. subdomains |
| Rejects | Shorts, live, embed, playlist and channel links |
| Source | Existing caption tracks, uploaded or auto-generated. Never audio |
| Billing | One event. Skipped videos are not charged, and there is no item fee |
| Timestamps | startTime is a string like "0:03"; dur is a string like "5.679" |
| Reliability | 99.95% of 4,087 runs succeeded in the last 30 days |
One video, a batch, or a batch in one language
Three runs. The billing model makes the second one unusually forgiving, and the third is where the one real subtlety in this actor lives.
How do you get a YouTube transcript?
Give the watch link. If the video has captions you get them back as timestamped lines for a tenth of a cent; if it does not, you are not charged at all.
Input
{
"startUrls": ["https://www.youtube.com/watch?v=_AbFXuGDRTs"],
"lang": "en"
}1 × $0.001=$0.001
What comes back
The transcript row, One row per video, holding every caption line as an array.
{
"inputSource": "https://www.youtube.com/watch?v=_AbFXuGDRTs",
"type": "transcript",
"id": "_AbFXuGDRTs",
"transcript": [
{ "text": "Hidden in this mountain is a $1 billion", "dur": "5.679", "startTime": "0:00" },
{ "text": "nuclear bunker. We are 2,000 ft", "dur": "4.8", "startTime": "0:03" }
]
}No captions means no transcript, not a worse one
This actor reads caption tracks; it does not run speech-to-text. A video with captions turned off is skipped with error code C003 and costs nothing, but no amount of retrying will produce text for it. If you need audio transcribed, this is the wrong tool entirely.
How do you transcribe a list of YouTube videos?
Put every watch link in startUrls. A skipped video does not stop the run and does not appear on the bill, so a hundred links where sixty have captions costs six cents, not ten.
Input
{
"startUrls": [
"https://www.youtube.com/watch?v=_AbFXuGDRTs",
"https://youtu.be/dQw4w9WgXcQ"
],
"maxItems": 100
}100 × $0.001=$0.10
What comes back
The caption line, Both time fields are strings, neither is a number or a timestamp type.
{
"text": "nuclear bunker. We are 2,000 ft",
"dur": "4.8",
"startTime": "0:03"
}The text is not clean prose
The listing warns that line text can carry line breaks, zero-width spaces, speaker-change markers (>>) and tags such as [Music], and that auto-generated tracks may show [ __ ] where YouTube has hidden a word. Strip these before anything reads the output as sentences.
How do you get transcripts in a specific language?
Set lang. But it is a request, not a filter: if the video has no track in that language a different one comes back, billed as normal. The listing’s own advice is to check selected.languageCode and drop the rows that do not match.
Input
{
"startUrls": ["https://www.youtube.com/watch?v=_AbFXuGDRTs"],
"lang": "fr",
"maxItems": 100
}100 × $0.001=$0.10
What comes back
The language blocks, One of these two is trustworthy. The other is documented as not being.
{
"selected": {
"title": "English (auto-generated)",
"languageCode": "en"
},
"availableLanguages": [
{ "title": "Vietnamese", "languageCode": "vi" }
]
}availableLanguages does not list the video’s languages
The listing says so outright: it is a single entry from the transcript backend and "doesn’t list the video’s caption tracks, so don’t use it to find languages". In the sample above it reports Vietnamese for a video whose returned track is English. Request a language with lang and read selected.languageCode. That is the only reliable answer.
One row per video, one array inside it
The shape is flat and small. Almost all of the payload sits in the transcript array, whose length is the length of the video rather than anything you control.
3
The row
Top level
- inputSource, the link you sent
- type, always "transcript"
- id, the video ID
3
Each caption line
transcript[]
- text
- startTime, string, "0:03"
- dur, string, "5.679"
2
What came back
selected
- title, e.g. "English (auto-generated)"
- languageCode
- The field to trust
2
Do not rely on
availableLanguages
- title, languageCode
- One backend entry, not the video’s tracks
inputSource | The exact link you submitted, how a row maps back to your list |
|---|---|
type | Always "transcript" on these records |
id | The YouTube video ID |
transcript | Array of caption lines, in order. Its length tracks the video’s length |
transcript[].text | One caption line. May contain line breaks, >> speaker markers, [Music] and [ __ ] |
transcript[].startTime | A string like "0:03", minutes and seconds, not a number and not ISO 8601 |
transcript[].dur | Seconds as a string, e.g. "5.679". Parse before summing |
selected.title | The returned track’s name, written in the requested language |
selected.languageCode | What you actually got. The field to check after every run using lang |
availableLanguages | A single entry from the transcript backend. The listing states it does not list the video’s caption tracks, do not use it for discovery |
inputSource is more useful here than on most actors. Because each link is processed separately and a duplicate link returns a duplicate row, it is the field that lets you reconcile a returned set against the list you submitted and see exactly which links came back empty.
The limits the listing states plainly
This is the most carefully documented actor in the catalogue, it names its own failure code and warns against one of its own fields. All five below are its words, not inferences.
- 01
It reads captions, it does not transcribe
Uploaded or auto-generated tracks only. A video with captions turned off returns nothing, and speech-to-text is listed among the things this actor does not do.
- 02
Only watch and youtu.be links
youtube.com/watch, including m. and music. subdomains, and youtu.be. Shorts, live, embed, playlist and channel links are rejected, as are links from other sites.
- 03
Skipped videos do not stop the batch or cost anything
A video without captions is skipped with error code C003 and the rest of the run continues. You are charged only for transcripts actually extracted.
- 04
lang is a request, not a guarantee
If the video has no track in the language you asked for, a different track is returned, and billed. Check selected.languageCode and discard the mismatches yourself.
- 05
availableLanguages is not a language list
The listing says it carries a single entry from the transcript backend and does not enumerate the video’s caption tracks. Using it to decide what to request will mislead you.
Free accounts run in demo mode: 5 runs a month, 10 items each.
One event, and nothing else
The simplest billing in the catalogue. There is a dataset-item event, and the listing states plainly that this actor never charges it, so the transcript price is the entire cost, with no storage fee on top and nothing billed for a failure.
| Run | Method | Calculation | Cost |
|---|---|---|---|
| 1 transcript | One video | 1 × $0.001 | $0.001 |
| 10 transcripts | A small batch | 10 × $0.001 | $0.01 |
| 100 transcripts | A channel’s back catalogue | 100 × $0.001 | $0.10 |
| 1,000 transcripts | A corpus | 1,000 × $0.001 | $1.00 |
| 10,000 transcripts | A training set | 10,000 × $0.001 | $10.00 |
| 0 transcripts | 40 videos, none captioned | nothing charged | $0.00 |
The last row is the one that changes how you work. Because failures are free, there is no penalty for throwing a whole channel’s links at it and keeping whatever comes back, which is not true of any other actor in the family.
Who needs the words out of the video?
Transcripts are the cheapest route from video to text, and text is what every search index, language model and analyst actually works on. The constraint is that the captions have to exist first.
LLM and RAG pipelines
A batch
Turn a channel into a retrievable corpus, with timestamps kept so an answer can cite the moment
$1.00 per 1,000 videos
Search and indexing
A batch, then re-run for new uploads
Make a video library searchable by what was said rather than by what the title claims
$0.10 per 100 videos
Media monitoring
A batch, scheduled
Catch mentions inside long videos that no title or description would ever surface
$0.01 per 10 videos
Accessibility and localisation
One language
Collect the caption tracks that exist, checking selected.languageCode to see which languages a library really covers
$0.10 per 100 checks
Research corpora
A batch, links from a playlist
Build a labelled text set from a curated list, pairing this with the Playlist Scraper to get the links
$10.00 per 10,000
From links to clean text
The actor is apidojo/youtube-transcript-scraper. Two things belong in every integration: check selected.languageCode, and strip the caption markup before anything treats the text as prose.
import re
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
MARKUP = re.compile(r"\[[^\]]*\]|>>|\u200b") # [Music], [ __ ], >>, zero-width
def clean(line):
return re.sub(r"\s+", " ", MARKUP.sub(" ", line)).strip()
run = client.actor("apidojo/youtube-transcript-scraper").call(run_input={
"startUrls": [
"https://www.youtube.com/watch?v=_AbFXuGDRTs",
"https://youtu.be/dQw4w9WgXcQ",
],
"lang": "en",
})
for row in client.dataset(run.default_dataset_id).iterate_items():
# lang is a request, not a filter: a miss returns another track, billed.
got = row["selected"]["languageCode"]
if got != "en":
print(f"skipping {row['id']}: got {got}")
continue
text = " ".join(clean(l["text"]) for l in row["transcript"])
seconds = sum(float(l["dur"]) for l in row["transcript"]) # dur is a string
print(f"{row['id']} {seconds/60:.1f} min {len(text):,} chars")import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('apidojo/youtube-transcript-scraper').call({
startUrls: ['https://www.youtube.com/watch?v=_AbFXuGDRTs'],
lang: 'en',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
// Videos without captions are skipped and never billed, so a short result
// set is normal. inputSource reconciles what came back against what you sent.
const returned = new Set(items.map((r) => r.inputSource));
console.log(items.length, 'transcripts;', returned.size, 'of your links came back');
// startTime is "0:03", not a number, keep it for citations, parse for maths.
for (const row of items) {
const chunks = row.transcript.map((l) => ({
at: l.startTime,
seconds: Number(l.dur),
text: l.text.replace(/\[[^\]]*\]|>>/g, '').trim(),
}));
console.log(row.id, row.selected.languageCode, chunks.length, 'lines');
}curl -X POST \
"https://api.apify.com/v2/acts/apidojo~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":["https://www.youtube.com/watch?v=_AbFXuGDRTs"],"lang":"en"}'JSON is the sensible export, CSV will flatten the transcript array into something you cannot read back. customMapFunction is the tidy place to join the lines into one field if that is all you need.
YouTube Transcript Scraper FAQ
Does this transcribe YouTube audio?
No. It reads the caption tracks a video already has, uploaded or auto-generated. A video with captions turned off returns no transcript, and speech-to-text is explicitly not supported.
How much does a YouTube transcript cost?
$0.001 each, or $1 per 1,000. There is no dataset-item fee on top: the listing states this actor does not charge that event at all.
Am I charged for videos with no captions?
No. They are skipped with error code C003, the rest of the batch continues, and only transcripts actually extracted are billed.
Can I scrape transcripts from Shorts?
No. Shorts, live, embed, playlist and channel links are all rejected. Only youtube.com/watch links, including m. and music. subdomains, and youtu.be links are accepted.
How do I get a transcript in a particular language?
Set lang to a code such as en, es or pt-BR. If the video has no track in that language a different one is returned and still billed, so check selected.languageCode and drop mismatches.
Can I see which languages a video has captions in?
Not from availableLanguages. The listing says it carries a single entry from the transcript backend and does not list the video’s caption tracks. Request a language and read selected.languageCode instead.
Are the timestamps numbers?
No. startTime is a string like "0:03" and dur is a string like "5.679". Keep them as they are for citation and parse them when you need arithmetic.
How do I transcribe a whole playlist or channel?
Collect the watch links first with the Playlist Scraper or the Channel Videos Scraper, then pass them to startUrls here, the listing recommends exactly that pairing.
Where YouTube Transcript Scraper fits
Parent scraper
YouTube Scraper
Seven YouTube scrapers. This is the text one, and the only one billed on success alone and a video it cannot read costs nothing, which is unusual enough to change how you batch.
Sibling scrapers
Guides
- Cleaning caption markup out of transcript textComing soon
- Building a citable RAG corpus from timestampsComing soon
- Playlist to transcript: chaining two scrapersComing soon
Pull one transcript
A tenth of a cent, and nothing at all if the captions are missing.
startUrls: [one watch link] · lang: "en"
Apidojo is not affiliated with YouTube or Google LLC.