Corpora collects YouTube video, audio, subtitles and structured metadata for multimodal AI training — vision-language models, speech models and video LLMs. Choose 1–40 Gbps unmetered yt-dlp proxies to run your own pipeline, or let us deliver curated datasets straight to your cloud storage.
From raw video and audio streams to clean, structured metadata — collected at scale, deduplicated and organized for vision-language, speech and video-LLM training.
Full-resolution video (up to 4K/8K), audio-only tracks for speech and music models, and selected formats. Delivered as original containers or transcoded to your target codec.
Title, description, tags, channel, publish date, duration, view/like counts, category, language and more — normalized to JSONL or Parquet.
Auto-generated and manual captions in all available languages, with timestamps aligned to the video for speech and multimodal training.
Top-level comments, replies, like counts and reply threads, useful for conversational, sentiment and ranking datasets.
All thumbnail resolutions plus optional keyframe extraction at custom intervals for vision-language pretraining.
Collect by keyword, channel list, playlist, topic, language, region, upload window or engagement thresholds. You define the scope.
Both options are built for sustained, high-volume collection without rate limits or bandwidth caps.
Dedicated 1–40 Gbps unmetered proxy capacity for your own collection stack.
Tell us what you need. We collect, process and deliver directly to your storage.
Every record is validated and normalized. Custom fields and schemas are available on request.
| Field | Type | Description |
|---|---|---|
video_id | string | Unique YouTube video identifier |
title / description | string | Original title and full description text |
channel_id / channel_name | string | Uploader channel identifiers and subscriber count snapshot |
published_at | datetime | Upload timestamp (UTC, ISO 8601) |
duration_sec | integer | Video length in seconds |
view_count / like_count / comment_count | integer | Engagement metrics at collection time |
tags / category | array / string | Creator tags and YouTube category |
language | string | Detected primary language (BCP-47) |
formats | array | Available resolutions, codecs, bitrates and file sizes |
subtitles | object | Available caption tracks by language, manual vs. auto-generated |
thumbnails | array | Thumbnail URLs and local paths for all resolutions |
file_path / sha256 | string | Location of delivered media in your bucket and integrity checksum |
Share your target: keywords, channels, languages, regions, time range, formats and estimated volume.
We deliver a free sample batch with metadata so you can verify quality, schema and format before committing.
Collection runs on dedicated 1–40 Gbps capacity, with progress dashboards and daily reports.
Data lands in your cloud storage with manifests and checksums. Incremental updates available on schedule.
Pay for bandwidth or pay for delivered data — whichever fits your pipeline. Volume discounts available on both.
Unmetered traffic. Scale from 1 Gbps to 40 Gbps.
Priced by delivered volume and processing scope. Contact us for a quote.
All prices in USD. Custom enterprise agreements, annual contracts and hybrid plans (proxies + delivery) are available — contact sales for details.
Yes. We collect full-resolution video, audio-only tracks, subtitles, thumbnails and structured metadata from YouTube at petabyte scale and deliver them to your cloud storage, or provide 1–40 Gbps unmetered proxies so you can run your own pipeline.
Yes. Our HTTP(S) and SOCKS5 proxies are fully compatible with yt-dlp, youtube-dl and custom crawlers, with unmetered traffic and rotating IP pools.
Video streams up to 4K/8K, audio tracks, manual and auto-generated captions with timestamps, thumbnails, extracted keyframes, comments and normalized JSONL/Parquet metadata — suitable for vision-language, speech and video-LLM training.
Proxies are $2,000 per Gbps per month with unmetered traffic and volume discounts from 5 Gbps. Managed dataset delivery is priced per delivered petabyte; a free sample batch is included.
800 TB per month is about 2.5 Gbps at 100% utilization. Accounting for protocol overhead and real-world utilization, we recommend 4–5 Gbps of dedicated capacity.
Share your target scope and volume. We'll respond with a sample plan, timeline and quote — typically within one business day.