Chapter 16

Claude Can Watch YouTube Now

Quick Install Guide for the claude-watch Skill

A free open-source Claude skill just shipped that does something most "video-to-text" tools fake: it actually watches the video. Reads every slide on screen. OCRs the code. Matches each frame to its timestamp. Writes you a structured markdown file you can re-read forever. MIT license. 8 stars in 2 days. Three install surfaces: Claude Code, claude.ai, and Codex. This guide gets you from zero to your first run in 5 minutes.

What It Actually Does

You paste a YouTube URL (or local video file). Walk away. Come back to a single markdown file with:

  • TLDR — 3-4 sentence synthesis
  • Key Concepts — bulleted with timestamps
  • Notes — one section per scene with embedded screenshot, on-screen text, what was said, and Claude's synthesis
  • Code & Commands — every code-on-screen frame transcribed into a runnable fenced block
  • Diagrams Referenced + Open Questions Files save to ~/claude-watch/library/<slug>/. Re-running the same URL = cache hit. No re-download. No re-transcribe. Free re-reads forever.

Why It's Better Than Uniform Sampling

Most video-to-text tools sample frames every N seconds. That wastes budget on a lecture where one slide stays on screen for 5 minutes (you get 1 frame for that whole stretch) and over-samples a fast-cut intro (10 nearly-identical frames). This skill is scene-aware. It uses ffmpeg to detect actual scene changes. Then it inserts a coverage-floor frame every 45 seconds across long static gaps so a 5-minute slide still gets ~7 frames captured, not 1. Result: a 30-minute lecture comes back with 30-50 well-chosen frames instead of 80 noisy ones.

Install (3 Surfaces)

Pick the one that matches where you use Claude.

/plugin marketplace add devinilabs/claude-watch
/plugin install claude-watch@claude-watch

That's it. The skill is now available as /claude-watch <url>.

claude.ai (web)

  1. Go to the GitHub releases page → download claude-watch.skill
  2. In claude.ai, open Settings → Capabilities → Skills
  3. Click + and upload the .skill file

Codex

git clone https://github.com/devinilabs/claude-watch ~/.codex/skills/claude-watch

Bring Your Own Keys (BYOK)

The skill is free to install. Running it costs almost nothing if videos have captions, pennies if they don't.

NeedCost
Download + native captionsFree (yt-dlp • ffmpeg)
Whisper fallback (preferred)Groq whisper-large-v3 — cheap, fast
Whisper fallback (alt)OpenAI whisper-1
Disable Whisper--no-whisper flag (frames-only)
Keys go in ~/.config/claude-watch/.env (mode 0600). I use Groq because their Whisper is 5-10x cheaper than OpenAI's and just as accurate.
mkdir -p ~/.config/claude-watch
cat > ~/.config/claude-watch/.env <<EOF
GROQ_API_KEY=gsk_xxx
EOF
chmod 600 ~/.config/claude-watch/.env

Your First Run

Pick a tutorial you've been meaning to watch. Paste this:

/claude-watch https://youtu.be/<your-video-id> backprop intuition

The second argument is the focus topic. It's optional but helps Claude prioritize what matters in the synthesis. For long lectures, narrow the range:

/claude-watch https://youtu.be/<long-video> --start 5:00 --end 25:00

For slides with tiny code text, bump the resolution:

/claude-watch <url> --resolution 1024

The 3 Commands I Run Daily

  1. Conference talks I missed live
/claude-watch <url> the actual demo

Forces Claude to focus on the working code, not the speaker bio.

  1. Tutorial videos for skills I'm learning
/claude-watch <url> --resolution 1024 step-by-step setup

Higher resolution catches every keystroke in the editor.

  1. Founder interviews I want to study
/claude-watch <url> --no-whisper key claims and timestamps

Frames-only mode when I just want the visual progression, not the talking head.

The Layer Most Operators Skip

Most people use this once, get a great notes file, never run it again. They treat it like a one-shot extractor. The move is to treat the library as a real KB.

~/claude-watch/library/
├── 2026-05-04-claude-code-walkthrough-a3b1/
├── 2026-05-05-anthropic-claude-skills-deep-dive-c2d4/
├── 2026-05-06-tldr-on-mcp-7e8f/

Every notes.md is just text. Which means:

  • Other Claude skills can Read them
  • Your agent can search across them via grep
  • You can pipe them into a vector DB for semantic search later Video input is a slow channel. But once it's converted to text + screenshots in a folder you own, it's the same shape as everything else your agents read. That's where it compounds.

What Stays Human

This skill turns video into notes. It does not turn notes into understanding. The parts you still do yourself:

  • Decide which videos are worth the run (most aren't)
  • Read the TLDR and the Open Questions section
  • Run any code blocks that look interesting
  • Connect what you just learned to a project you're working on If you skip the read-and-connect part, you've just built a graveyard of summaries you'll never look at again.

Where This Goes Next

The pattern under this skill — agents that write to a persistent library you can re-read — is what separates Claude-as-chat-window from Claude-as-operating-system. Video notes are one stream. Calls become briefs. Articles become KB entries. LinkedIn posts get archived as voice samples. They all flow into the same library. Future agents read everything you've ever produced and stay in your voice.

Built by Joon at AI Topia. joon@getaitopia.io YouTube @joonahn_ai · LinkedIn

Want this set up for your business?

We turn these workflows into working marketing and sales systems.

Book a call