AI Transcription
Every completed recording is transcribed automatically — word-level timestamps, three output formats, language auto-detection.
AI Transcription
VIDTREO AI transcribes every completed recording automatically. There is no transcription API to call and no audio to extract — when a video finishes uploading in an environment with transcription enabled, the pipeline runs on its own and notifies your systems via webhooks when the transcript is ready.
Recording completes
→ Audio extracted from the video
→ Transcription with word-level timestamps
→ Language auto-detected
→ Transcript stored in three formats (VTT, JSON, TXT)
→ transcriptions.transcription.completed webhook firesEnabling Transcription
Transcription is configured per environment:
Sign in to the VIDTREO Dashboard
Open your environment and go to Settings → General
Turn on AI Transcription
From that moment, every recording that completes in the environment is transcribed automatically. Videos uploaded while the toggle was off are not processed retroactively.
Language is detected automatically — there is no language parameter. Users can record in Spanish, Portuguese, English, or any other spoken language and the transcript comes back in that language.
Output Formats
Each transcription produces three files:
| Format | Content type | Built for |
|---|---|---|
vtt | text/vtt | WebVTT captions — drop straight into a <track> element for accessible playback |
json | application/json | Word-level segments with start/end times — search, jump-to-moment UIs, feeding LLMs |
txt | text/plain | Plain transcript — full-text indexing, summaries, archives |
The JSON format is an array of word-level segments:
[
{ "word": "welcome", "start": 1.02, "end": 1.38 },
{ "word": "to", "start": 1.38, "end": 1.5 },
{ "word": "the", "start": 1.5, "end": 1.62 }
]Transcription Lifecycle
| Status | Meaning |
|---|---|
PENDING | Queued, waiting to be processed |
PROCESSING | Audio extraction and transcription in progress |
COMPLETED | Transcript files are available |
FAILED | Processing failed — can be reprocessed from the dashboard |
SKIPPED | The video has no audio track — nothing to transcribe |
A video recorded without a microphone (for example, a muted screen recording) results in SKIPPED, not FAILED. Handle both statuses if your workflow expects a transcript for every video.
Fetching a Transcript via API
Retrieve the transcription for any video with your API key (requires the get_video permission):
curl https://core.vidtreo.com/api/v1/videos/{videoId}/transcription \
-H "Authorization: Bearer vt_live_xxxxxxxxxxxxxxxxxxxx"Response:
{
"id": "b7e2...",
"videoId": "vid_...",
"environmentId": "env_...",
"status": "COMPLETED",
"language": "en",
"wordCount": 812,
"duration": 294.5,
"files": {
"vtt": { "url": "https://...", "expiresAt": "...", "contentType": "text/vtt" },
"json": { "url": "https://...", "expiresAt": "...", "contentType": "application/json" },
"txt": { "url": "https://...", "expiresAt": "...", "contentType": "text/plain" }
},
"createdAt": "2026-08-03T14:17:00.000Z",
"completedAt": "2026-08-03T14:22:07.000Z"
}files is null until status is COMPLETED.
File URLs are presigned and expire 5 minutes after they are generated. Fetch the content when you receive them, store what you need on your side, and request fresh URLs from this endpoint whenever you need the files again. Never persist the URLs themselves.
Getting Notified
Polling this endpoint works, but the intended integration is event-driven: subscribe a webhook endpoint to the transcriptions.transcription.completed event and receive the transcript files — plus your own metadata — the moment processing finishes.
See Webhooks for the payload shape, signature verification, and delivery guarantees.