For: every JBNX agent that needs the words from a YouTube video. Slug: youtube-transcripts Page: https://projects.jbnx.io/framework/youtube-transcripts Do not re-learn this from scratch. Follow the ladder. Stop when you have the captions.
Sister framework (what to do with a strategy transcript): FrameWork Youtube Viral Videos.
A timestamped caption file covering the full runtime, plus the real title, channel, length, and chapter list. Auto-captions are ASR: names will be wrong. Keep them, note the caveats, do not "fix" quotes unless you watched that second.
Cloud / datacenter IPs are treated as bots. The watch page and InnerTube player often return LOGIN_REQUIRED / "Sign in to confirm you're not a bot" with empty captions. That is not "no captions exist."
Do this first:
curl -sS "https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=VIDEO_ID&format=json"
You get title, author_name, author_url, thumbnail. Length is not in oembed.
Optional InnerTube that often returns metadata only (title, lengthSeconds, description + chapters) even when playability is UNPLAYABLE:
POST https://www.youtube.com/youtubei/v1/player
clientName: ANDROID_TESTSUITE / clientVersion: 1.9
Read videoDetails.title, lengthSeconds, shortDescription (chapters live at the bottom of the description on long Open Residency / podcast uploads).
Never trust a web-search snippet for which video an 11-character id is. Similar Open Residency titles collide (Kallaway vs Paddy Galloway). oembed / videoDetails wins.
Skip these once you have seen bot-block or empty timedtext. They burn the hour:
| Path | Typical failure |
|---|---|
youtube-transcript-api | RequestBlocked (cloud IP) |
yt-dlp without cookies + JS runtime | Sign in to confirm you're not a bot |
InnerTube ANDROID / IOS / TVHTML5 / WEB_EMBEDDED | LOGIN_REQUIRED or 400 |
https://www.youtube.com/api/timedtext?v=…&lang=en | empty 200 |
| Invidious / Piped public instances | 401/403/502 or shutdown |
| Jina / many "transcript generator" sites | Cloudflare challenge |
Watch-page WebFetch | timeout or bot interstitial |
A second Cloud Agent will hit the same IP reputation. Do not spawn one to "verify."
Open https://www.youtube.com/watch?v=VIDEO_ID in a real browser (computer-use / headed Chrome). Accept cookies. Do not Google-sign-in.
YouTube glues the visible stamp to an accessibility label:
0:000 secondsWithin a day of changing…
0:1414 secondsVideos that have 10, 50…
1:041 minute, 4 secondsYouTube strategy.
2:44:252 hours, 44 minutes, 25 secondsbigger. So, I think…
Chapter 12: Ten Titles, Three Thumbnails, Every Video
That is:
0:00 + 0 seconds + text0:14 + 14 seconds + text1:04 + 1 minute, 4 seconds + text2:44:25 + 2 hours, 44 minutes, 25 seconds + text2:002 minutesYeah. → 2:00 + 2 minutes + textParser (copy this, do not reinvent):
import re
a11y = re.compile(
r'^(?P<ts>\d{1,2}:\d{2}(?::\d{2})?)'
r'(?:'
r'(?:(?P<h>\d+) hours?, )?'
r'(?:(?P<m>\d+) minutes?, )?'
r'(?P<s>\d+) seconds?'
r'|'
r'(?:(?P<h2>\d+) hours?, )?'
r'(?P<m2>\d+) minutes?'
r')'
r'(?P<text>.*)$'
)
ch_re = re.compile(r'^Chapter \d+:\s*(.+)$')
def norm_ts(ts: str) -> str:
parts = [int(x) for x in ts.split(":")]
if len(parts) == 2:
mm, ss = parts
return f"0:{mm:02d}:{ss:02d}" # first hour is M:SS
h, mm, ss = parts
return f"{h}:{mm:02d}:{ss:02d}"
Chapter N: lines as headers.New, Nmo ago, a bare H:MM:SS duration on its own line, then another video title).lengthSeconds; spoken-word count roughly duration_min * 130–160.[music] is ASR for score under speech. Strip for reading; keep in the archival file.Many long YouTube interviews ship the same audio to Apple/Spotify the same day.
curl -sS "https://itunes.apple.com/search?term=VIDEO_TITLE&entity=podcastEpisode&limit=5"
You get feedUrl, episodeUrl / previewUrl (mp3). Use this when you need audio (Whisper) because captions never loaded — not as a substitute for captions when the panel worked.
Episode pages (example: https://openresidency.com/<guest-kebab>) give chapters even when YouTube is blocked.
Export cookies from a human browser (yt-dlp --cookies) only if the operator offered them. Do not burn a Google account. Whisper on the podcast mp3 if captions truly do not exist (rare on long English uploads).
Write three artifacts, then answer in chat:
*-raw.txt) — clipboard, untouched.*-final.txt) — [H:MM:SS] text plus ## Chapter headings.Do not commit a 40k-word transcript into the JBNX repo or public/framework/. The framework stays here; the transcript stays in chat / /tmp / artifacts.
| Field | Value |
|---|---|
| URL | https://www.youtube.com/watch?v=Z2uoA3bhJT0 |
| oembed title | 2-Hour Youtube Masterclass From The World's Highest-Paid Strategist |
| Channel | Open Residency |
| Guest | Paddy Galloway (ASR: "Patty") |
| Length | 9908 s = 2:45:08 |
| Method that worked | Browser Show transcript → a11y parse |
| Captions | 1346 lines, 24 chapters, 0:00:00 → 2:45:01, ~38.8k words |
| Search trap | Web search matched a different Open Residency cut (Kallaway, VcqQmrGqthg). oembed prevented mixing them. |
/go/call as proof of anything.--done + release.public/framework/. New frameworks go to portfolio.documents (this page). Public agent-readable slugs must also sit in PUBLIC_DOCUMENTS in projects-portal/owner-auth.mjs.When the browser runtime advertises tab.content.exportYouTubeTranscript(), use that supported export on the identified watch page before manual panel copying. Verify the exported video ID, caption language, timestamp coverage and runtime. This worked for u_yvc7NTYvI (Dubibubi, 16:20; English ASR 0:00–16:17) after HTTP timedtext returned empty responses. Empty captions from one access path do not prove the video has no transcript.
Separate transcript-confirmed statements from independently verified technical facts. Preserve the difference between headline maxima and the actual demo result; do not validate an unseen screenshot from an AI description. See the completed rule review. Never publish the full copyrighted transcript as a framework.