AI Training Data Collection

YouTube Video & Audio Datasets for AI Training, at Petabyte Scale

Corpora collects YouTube video, audio, subtitles and structured metadata for multimodal AI training — vision-language models, speech models and video LLMs. Choose 1–40 Gbps unmetered yt-dlp proxies to run your own pipeline, or let us deliver curated datasets straight to your cloud storage.

YouTube video and audio dataset collection pipeline for multimodal AI training
1–40 GbpsDedicated bandwidth per client, unmetered traffic
PB-scaleVideo, audio and metadata delivered to your bucket
99.9%Proxy uptime with automatic failover
What We Collect

Multimodal training data from YouTube — video, audio, text and metadata

From raw video and audio streams to clean, structured metadata — collected at scale, deduplicated and organized for vision-language, speech and video-LLM training.

YouTube Video & Audio Download

Full-resolution video (up to 4K/8K), audio-only tracks for speech and music models, and selected formats. Delivered as original containers or transcoded to your target codec.

{ }

Structured Metadata

Title, description, tags, channel, publish date, duration, view/like counts, category, language and more — normalized to JSONL or Parquet.

CC

Subtitles & Transcripts

Auto-generated and manual captions in all available languages, with timestamps aligned to the video for speech and multimodal training.

💬

Comments & Engagement

Top-level comments, replies, like counts and reply threads, useful for conversational, sentiment and ranking datasets.

🖼

Thumbnails & Frames

All thumbnail resolutions plus optional keyframe extraction at custom intervals for vision-language pretraining.

Custom Targeting

Collect by keyword, channel list, playlist, topic, language, region, upload window or engagement thresholds. You define the scope.

Two Ways to Work With Us

Bring your own pipeline, or let us handle delivery

Both options are built for sustained, high-volume collection without rate limits or bandwidth caps.

Self-managed

High-Bandwidth Proxies

Dedicated 1–40 Gbps unmetered proxy capacity for your own collection stack.

  • Unmetered traffic — no per-GB charges, ever
  • Scale from 1 Gbps to 40 Gbps per client
  • HTTP(S) and SOCKS5, compatible with yt-dlp and custom crawlers
  • Rotating IP pools with automatic health checks and failover
  • Multiple regions for geo-specific content
  • 24/7 technical support and dedicated account manager
See Proxy Pricing
Fully managed

Cloud Storage Delivery

Tell us what you need. We collect, process and deliver directly to your storage.

  • Delivered to your S3, GCS, Azure Blob or private object storage
  • Video, audio, subtitles and metadata in your preferred layout
  • Deduplication, integrity checksums and manifest files included
  • Optional transcoding, frame extraction and format conversion
  • Incremental updates for continuously growing datasets
  • Priced per delivered PB — no infrastructure on your side
Request a Quote
Metadata Schema

Clean, consistent fields for every video

Every record is validated and normalized. Custom fields and schemas are available on request.

FieldTypeDescription
video_idstringUnique YouTube video identifier
title / descriptionstringOriginal title and full description text
channel_id / channel_namestringUploader channel identifiers and subscriber count snapshot
published_atdatetimeUpload timestamp (UTC, ISO 8601)
duration_secintegerVideo length in seconds
view_count / like_count / comment_countintegerEngagement metrics at collection time
tags / categoryarray / stringCreator tags and YouTube category
languagestringDetected primary language (BCP-47)
formatsarrayAvailable resolutions, codecs, bitrates and file sizes
subtitlesobjectAvailable caption tracks by language, manual vs. auto-generated
thumbnailsarrayThumbnail URLs and local paths for all resolutions
file_path / sha256stringLocation of delivered media in your bucket and integrity checksum
How It Works

From requirements to delivered dataset

STEP 01

Define Scope

Share your target: keywords, channels, languages, regions, time range, formats and estimated volume.

STEP 02

Sample & Validate

We deliver a free sample batch with metadata so you can verify quality, schema and format before committing.

STEP 03

Collect at Scale

Collection runs on dedicated 1–40 Gbps capacity, with progress dashboards and daily reports.

STEP 04

Deliver & Verify

Data lands in your cloud storage with manifests and checksums. Incremental updates available on schedule.

Pricing

Simple, volume-friendly pricing

Pay for bandwidth or pay for delivered data — whichever fits your pipeline. Volume discounts available on both.

Managed Dataset Delivery
Per PBdelivered

Priced by delivered volume and processing scope. Contact us for a quote.

  • Video, audio, subtitles and metadata included
  • Delivered to S3 / GCS / Azure / private storage
  • Deduplication, checksums and manifests
  • Optional transcoding and frame extraction
Free sample batchIncluded
Multi-PB projectsTiered discount
Ongoing incremental deliveryCustom terms
Request a Quote

All prices in USD. Custom enterprise agreements, annual contracts and hybrid plans (proxies + delivery) are available — contact sales for details.

FAQ

Frequently asked questions

Can you provide YouTube video and audio datasets for AI training?

Yes. We collect full-resolution video, audio-only tracks, subtitles, thumbnails and structured metadata from YouTube at petabyte scale and deliver them to your cloud storage, or provide 1–40 Gbps unmetered proxies so you can run your own pipeline.

Do your proxies work with yt-dlp?

Yes. Our HTTP(S) and SOCKS5 proxies are fully compatible with yt-dlp, youtube-dl and custom crawlers, with unmetered traffic and rotating IP pools.

What multimodal data types do you deliver?

Video streams up to 4K/8K, audio tracks, manual and auto-generated captions with timestamps, thumbnails, extracted keyframes, comments and normalized JSONL/Parquet metadata — suitable for vision-language, speech and video-LLM training.

How is pricing structured?

Proxies are $2,000 per Gbps per month with unmetered traffic and volume discounts from 5 Gbps. Managed dataset delivery is priced per delivered petabyte; a free sample batch is included.

How much bandwidth do I need to download 800 TB per month?

800 TB per month is about 2.5 Gbps at 100% utilization. Accounting for protocol overhead and real-world utilization, we recommend 4–5 Gbps of dedicated capacity.

Contact Sales

Tell us about your dataset

Share your target scope and volume. We'll respond with a sample plan, timeline and quote — typically within one business day.