# YouTube Transcript Scraper: captions, not speech-to-text

YouTube Transcript Scraper returns the caption track of a public YouTube video as structured lines, each with its text, its start time and its duration, reading the captions the video already has rather than transcribing its audio.

[Run this scraper free](https://apify.com/apidojo/youtube-transcript-scraper)

[See the three jobs](#jobs)

apidojo/youtube-transcript-scraper

**Transcripts extracted**

Estimated run cost

$0.00

Free Apify plan: 5 runs a month, 10 items a run. No credit card.

**$0.001** per transcript extracted

**$0** for a video with no captions

**$1** per 1,000 transcripts

**99.95%** run success, last 30 days

Read on 27 September 2026.

## Four inputs, and a list of links it will refuse

The input surface is small and the listing is precise about its edges, which link forms work, which are rejected, and what happens to the ones in between.

-   `startUrls`

    $0.001per transcript

    Only if it succeeds

    One video link per entry, as a plain string or as { "url": "…" }. Each is processed separately, so a duplicate link returns a duplicate row.

    `https://www.youtube.com/watch?v=_AbFXuGDRTs`

-   A miss does not fail. It returns another track

    `lang`

    optionalno charge

    Check what came back

    An optional caption language code, en, es, de, tr, pt-BR. Leave it empty for the video’s default track.

    `lang: "fr"`

-   `maxItems`

    ÷ $0.001caps the bill

    Run-wide

    Stops the run saving once this many rows exist. Leave it empty to process every link you supplied.

    `100`

-   `Unsupported links`

    rejectednot charged

    watch and youtu.be only

    Shorts, live, embed, playlist and channel URLs are rejected outright, as are links from other sites.

    `/shorts/… → rejected`

Starting from a playlist or a channel? The listing points you at the Playlist and Channel Videos scrapers to collect the watch links first, then bring them here, which is exactly the pairing those two pages describe from the other end.

**Apify actor:** apidojo/youtube-transcript-scraper

**Inputs:** startUrls · lang · maxItems · customMapFunction

**Accepts:** youtube.com/watch and youtu.be links, including m. and music. subdomains

**Rejects:** Shorts, live, embed, playlist and channel links

**Source:** Existing caption tracks, uploaded or auto-generated. Never audio

**Billing:** One event. Skipped videos are not charged, and there is no item fee

**Timestamps:** startTime is a string like "0:03"; dur is a string like "5.679"

**Reliability:** 99.95% of 4,087 runs succeeded in the last 30 days

## One video, a batch, or a batch in one language

Three runs. The billing model makes the second one unusually forgiving, and the third is where the one real subtlety in this actor lives.

**01 One video $0.001, on success only**

**02 A batch Only successes are billed**

**03 One language Mismatches still cost $0.001**

### How do you get a YouTube transcript?

Give the watch link. If the video has captions you get them back as timestamped lines for a tenth of a cent; if it does not, you are not charged at all.

Input

```
{
  "startUrls": ["https://www.youtube.com/watch?v=_AbFXuGDRTs"],
  "lang": "en"
}
```

1 × $0.001 \= **$0.001**

What comes back

The **transcript row**, One row per video, holding every caption line as an array.

```
{
  "inputSource": "https://www.youtube.com/watch?v=_AbFXuGDRTs",
  "type": "transcript",
  "id": "_AbFXuGDRTs",
  "transcript": [
    { "text": "Hidden in this mountain is a $1 billion", "dur": "5.679", "startTime": "0:00" },
    { "text": "nuclear bunker. We are 2,000 ft", "dur": "4.8", "startTime": "0:03" }
  ]
}
```

No captions means no transcript, not a worse one

This actor reads caption tracks; it does not run speech-to-text. A video with captions turned off is skipped with error code C003 and costs nothing, but no amount of retrying will produce text for it. If you need audio transcribed, this is the wrong tool entirely.

### How do you transcribe a list of YouTube videos?

Put every watch link in startUrls. A skipped video does not stop the run and does not appear on the bill, so a hundred links where sixty have captions costs six cents, not ten.

Input

```
{
  "startUrls": [
    "https://www.youtube.com/watch?v=_AbFXuGDRTs",
    "https://youtu.be/dQw4w9WgXcQ"
  ],
  "maxItems": 100
}
```

100 × $0.001 \= **$0.10**

What comes back

The **caption line**, Both time fields are strings, neither is a number or a timestamp type.

```
{
  "text": "nuclear bunker. We are 2,000 ft",
  "dur": "4.8",
  "startTime": "0:03"
}
```

The text is not clean prose

The listing warns that line text can carry line breaks, zero-width spaces, speaker-change markers (\>\>) and tags such as \[Music\], and that auto-generated tracks may show \[ \_\_ \] where YouTube has hidden a word. Strip these before anything reads the output as sentences.

### How do you get transcripts in a specific language?

Set lang. But it is a request, not a filter: if the video has no track in that language a different one comes back, billed as normal. The listing’s own advice is to check selected.languageCode and drop the rows that do not match.

Input

```
{
  "startUrls": ["https://www.youtube.com/watch?v=_AbFXuGDRTs"],
  "lang": "fr",
  "maxItems": 100
}
```

100 × $0.001 \= **$0.10**

What comes back

The **language blocks**, One of these two is trustworthy. The other is documented as not being.

```
{
  "selected": {
    "title": "English (auto-generated)",
    "languageCode": "en"
  },
  "availableLanguages": [
    { "title": "Vietnamese", "languageCode": "vi" }
  ]
}
```

availableLanguages does not list the video’s languages

The listing says so outright: it is a single entry from the transcript backend and "doesn’t list the video’s caption tracks, so don’t use it to find languages". In the sample above it reports Vietnamese for a video whose returned track is English. Request a language with lang and read selected.languageCode. That is the only reliable answer.

## One row per video, one array inside it

The shape is flat and small. Almost all of the payload sits in the transcript array, whose length is the length of the video rather than anything you control.

-   3

    The row

    Top level

    -   inputSource, the link you sent
    -   type, always "transcript"
    -   id, the video ID
-   3

    Each caption line

    transcript\[\]

    -   text
    -   startTime, string, "0:03"
    -   dur, string, "5.679"
-   2

    What came back

    selected

    -   title, e.g. "English (auto-generated)"
    -   languageCode
    -   The field to trust
-   2

    Do not rely on

    availableLanguages

    -   title, languageCode
    -   One backend entry, not the video’s tracks

**inputSource:** The exact link you submitted, how a row maps back to your list

**type:** Always "transcript" on these records

**id:** The YouTube video ID

**transcript:** Array of caption lines, in order. Its length tracks the video’s length

**transcript[].text:** One caption line. May contain line breaks, >> speaker markers, [Music] and [ __ ]

**transcript[].startTime:** A string like "0:03", minutes and seconds, not a number and not ISO 8601

**transcript[].dur:** Seconds as a string, e.g. "5.679". Parse before summing

**selected.title:** The returned track’s name, written in the requested language

**selected.languageCode:** What you actually got. The field to check after every run using lang

**availableLanguages:** A single entry from the transcript backend. The listing states it does not list the video’s caption tracks, do not use it for discovery

inputSource is more useful here than on most actors. Because each link is processed separately and a duplicate link returns a duplicate row, it is the field that lets you reconcile a returned set against the list you submitted and see exactly which links came back empty.

## The limits the listing states plainly

This is the most carefully documented actor in the catalogue, it names its own failure code and warns against one of its own fields. All five below are its words, not inferences.

-   01

    It reads captions, it does not transcribe

    Uploaded or auto-generated tracks only. A video with captions turned off returns nothing, and speech-to-text is listed among the things this actor does not do.

-   02

    Only watch and youtu.be links

    youtube.com/watch, including m. and music. subdomains, and youtu.be. Shorts, live, embed, playlist and channel links are rejected, as are links from other sites.

-   03

    Skipped videos do not stop the batch or cost anything

    A video without captions is skipped with error code C003 and the rest of the run continues. You are charged only for transcripts actually extracted.

-   04

    lang is a request, not a guarantee

    If the video has no track in the language you asked for, a different track is returned, and billed. Check selected.languageCode and discard the mismatches yourself.

-   05

    availableLanguages is not a language list

    The listing says it carries a single entry from the transcript backend and does not enumerate the video’s caption tracks. Using it to decide what to request will mislead you.

Free accounts run in demo mode: 5 runs a month, 10 items each.

## One event, and nothing else

The simplest billing in the catalogue. There is a dataset-item event, and the listing states plainly that this actor never charges it, so the transcript price is the entire cost, with no storage fee on top and nothing billed for a failure.

| Run | Method | Calculation | Cost |
| --- | --- | --- | --- |
| 1 transcript | One video | 1 × $0.001 | $0.001 |
| 10 transcripts | A small batch | 10 × $0.001 | $0.01 |
| 100 transcripts | A channel’s back catalogue | 100 × $0.001 | $0.10 |
| 1,000 transcripts | A corpus | 1,000 × $0.001 | $1.00 |
| 10,000 transcripts | A training set | 10,000 × $0.001 | $10.00 |
| 0 transcripts | 40 videos, none captioned | nothing charged | $0.00 |

The last row is the one that changes how you work. Because failures are free, there is no penalty for throwing a whole channel’s links at it and keeping whatever comes back, which is not true of any other actor in the family.

## Who needs the words out of the video?

Transcripts are the cheapest route from video to text, and text is what every search index, language model and analyst actually works on. The constraint is that the captions have to exist first.

-   LLM and RAG pipelines

    A batch

    Turn a channel into a retrievable corpus, with timestamps kept so an answer can cite the moment

    $1.00 per 1,000 videos

-   Search and indexing

    A batch, then re-run for new uploads

    Make a video library searchable by what was said rather than by what the title claims

    $0.10 per 100 videos

-   Media monitoring

    A batch, scheduled

    Catch mentions inside long videos that no title or description would ever surface

    $0.01 per 10 videos

-   Accessibility and localisation

    One language

    Collect the caption tracks that exist, checking selected.languageCode to see which languages a library really covers

    $0.10 per 100 checks

-   Research corpora

    A batch, links from a playlist

    Build a labelled text set from a curated list, pairing this with the Playlist Scraper to get the links

    $10.00 per 10,000

## From links to clean text

The actor is apidojo/youtube-transcript-scraper. Two things belong in every integration: check selected.languageCode, and strip the caption markup before anything treats the text as prose.

**Python**

**JavaScript**

**cURL**

```
import re
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

MARKUP = re.compile(r"\[[^\]]*\]|>>|\u200b")   # [Music], [ __ ], >>, zero-width

def clean(line):
    return re.sub(r"\s+", " ", MARKUP.sub(" ", line)).strip()

run = client.actor("apidojo/youtube-transcript-scraper").call(run_input={
    "startUrls": [
        "https://www.youtube.com/watch?v=_AbFXuGDRTs",
        "https://youtu.be/dQw4w9WgXcQ",
    ],
    "lang": "en",
})

for row in client.dataset(run.default_dataset_id).iterate_items():
    # lang is a request, not a filter: a miss returns another track, billed.
    got = row["selected"]["languageCode"]
    if got != "en":
        print(f"skipping {row['id']}: got {got}")
        continue

    text = " ".join(clean(l["text"]) for l in row["transcript"])
    seconds = sum(float(l["dur"]) for l in row["transcript"])   # dur is a string
    print(f"{row['id']}  {seconds/60:.1f} min  {len(text):,} chars")
```

```
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });

const run = await client.actor('apidojo/youtube-transcript-scraper').call({
  startUrls: ['https://www.youtube.com/watch?v=_AbFXuGDRTs'],
  lang: 'en',
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();

// Videos without captions are skipped and never billed, so a short result
// set is normal. inputSource reconciles what came back against what you sent.
const returned = new Set(items.map((r) => r.inputSource));
console.log(items.length, 'transcripts;', returned.size, 'of your links came back');

// startTime is "0:03", not a number, keep it for citations, parse for maths.
for (const row of items) {
  const chunks = row.transcript.map((l) => ({
    at: l.startTime,
    seconds: Number(l.dur),
    text: l.text.replace(/\[[^\]]*\]|>>/g, '').trim(),
  }));
  console.log(row.id, row.selected.languageCode, chunks.length, 'lines');
}
```

```
curl -X POST \
  "https://api.apify.com/v2/acts/apidojo~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":["https://www.youtube.com/watch?v=_AbFXuGDRTs"],"lang":"en"}'
```

JSON is the sensible export, CSV will flatten the transcript array into something you cannot read back. customMapFunction is the tidy place to join the lines into one field if that is all you need.

## YouTube Transcript Scraper FAQ

### Does this transcribe YouTube audio?

No. It reads the caption tracks a video already has, uploaded or auto-generated. A video with captions turned off returns no transcript, and speech-to-text is explicitly not supported.

### How much does a YouTube transcript cost?

$0.001 each, or $1 per 1,000. There is no dataset-item fee on top: the listing states this actor does not charge that event at all.

### Am I charged for videos with no captions?

No. They are skipped with error code C003, the rest of the batch continues, and only transcripts actually extracted are billed.

### Can I scrape transcripts from Shorts?

No. Shorts, live, embed, playlist and channel links are all rejected. Only youtube.com/watch links, including m. and music. subdomains, and youtu.be links are accepted.

### How do I get a transcript in a particular language?

Set lang to a code such as en, es or pt-BR. If the video has no track in that language a different one is returned and still billed, so check selected.languageCode and drop mismatches.

### Can I see which languages a video has captions in?

Not from availableLanguages. The listing says it carries a single entry from the transcript backend and does not list the video’s caption tracks. Request a language and read selected.languageCode instead.

### Are the timestamps numbers?

No. startTime is a string like "0:03" and dur is a string like "5.679". Keep them as they are for citation and parse them when you need arithmetic.

### How do I transcribe a whole playlist or channel?

Collect the watch links first with the Playlist Scraper or the Channel Videos Scraper, then pass them to startUrls here, the listing recommends exactly that pairing.

## Where YouTube Transcript Scraper fits

[Parent scraper YouTube Scraper Seven YouTube scrapers. This is the text one, and the only one billed on success alone and a video it cannot read costs nothing, which is unusual enough to change how you batch.](/scrapers/youtube-scraper/)

Sibling scrapers

-   [YouTube Playlist Scraper](/scrapers/youtube-playlist-scraper/)
-   [YouTube Channel Videos Scraper](/scrapers/youtube-channel-scraper/)
-   [YouTube Comments Scraper](/scrapers/youtube-comments-scraper/)

Guides

-   Cleaning caption markup out of transcript text*Coming soon*
-   Building a citable RAG corpus from timestamps*Coming soon*
-   Playlist to transcript: chaining two scrapers*Coming soon*

## Pull one transcript

A tenth of a cent, and nothing at all if the captions are missing.

`startUrls: [one watch link] · lang: "en"`

[Run this scraper free](https://apify.com/apidojo/youtube-transcript-scraper)

Apidojo is not affiliated with YouTube or Google LLC.
