API para TikTok: Transcrição de vídeo (versão completa)
Transcreva o áudio falado de um vídeo do TikTok com o speech-to-text próprio da AnyAPI: segmentos com tempo por frase, rótulos de locutor, idioma detectado e tempos por palavra opcionais, para vídeos em que o TikTok não publica faixa de legenda e para vídeos cuja faixa de legenda erra as palavras. A AnyAPI baixa o vídeo e processa o áudio com o MAI-Transcribe-2 em vez de ler qualquer coisa que o TikTok escreveu, e é por isso que lida com fala em outros idiomas além do inglês e ouve palavras que as legendas nativas entendem errado. A resposta também traz o ID do vídeo, a URL, o nome de usuário do criador e a imagem de capa. Ative hostVideo para também receber o MP4 em um link hospedado que toca sem depender da URL de CDN assinada e de curta duração do TikTok, que expira. Se você só quer o que o próprio TikTok publicou, e por um décimo do preço, use tiktok.video_transcript.
Teste
Faça sua primeira requisição
{
"data": {
"bytes": 42,
"durationSeconds": 12.5,
"expiresUtc": 12.5,
"hostedUrl": "https://example.com/page",
"id": "a1b2c3d4",
"language": "en",
"ownerUsername": "alex_rivera",
"segments": [
{
"endSeconds": 180,
"language": "en",
"speaker": "example",
"startSeconds": 180,
"text": "A short example description of this item.",
"words": [
{
"endSeconds": 180,
"startSeconds": 180,
"text": "A short example description of this item."
}
]
}
],
"source": "audio_asr",
"thumbnailUrl": "https://example.com/image.jpg",
"transcript": "example",
"url": "https://example.com/page"
},
"found": true,
"reason": "not_found"
}interface TiktokVideoTranscriptFullResponse {
data: {
bytes?: number;
durationSeconds?: number;
expiresUtc?: number;
hostedUrl?: string;
id?: string;
language?: string;
ownerUsername?: string;
segments?: {
endSeconds: number;
language?: string;
speaker?: string;
startSeconds: number;
text: string;
words?: {
endSeconds?: number;
startSeconds?: number;
text: string;
}[];
}[];
source: "audio_asr" | "transcript_unavailable";
thumbnailUrl?: string;
transcript: string;
url?: string;
} | null;
found: boolean;
reason?: "not_found";
}Referência completa de parâmetros e resposta - todos os campos, tipos e exemplos deste endpoint.
Referência
Requisição, resposta e preço
Última verificação em 2026-09-29 · disponibilidade e latência medidas em 30dcurl -X POST https://api.getanyapi.com/v1/run/tiktok.video_transcript_full \
-H "Authorization: Bearer $ANYAPI_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://www.tiktok.com/@thatdudecancook/video/7649086431641521421"}'| Campo | Tipo | Valor de exemplo |
|---|---|---|
| Corpo da requisição | ||
| url | string | "https://www.tiktok.com/@thatdudecancook/video/7649086431641521421"URL do vídeo do TikTok (ex.: "https://www.tiktok.com/@user/video/1234567890"). |
| hostVideo | boolean | Também armazena o vídeo e retorna um link de MP4 hospedado que toca sem a URL de CDN assinada do TikTok. Cobrado como adicional, além da transcrição. |
| wordTimestamps | boolean | Retorna tempos por palavra dentro de cada segmento. As palavras trazem apenas tempos; o reconhecedor pontua uma frase, não uma palavra, então não há confiança por palavra para informar. |
| Resposta | ||
| data | object | |
| data.bytes | integer | Size of the hosted MP4 in bytes. |
| data.durationSeconds | number | Video duration in seconds. |
| data.expiresUtc | number | When the hosted link stops working. UTC epoch timestamp in seconds (Unix time). Multiply by 1000 for a JS Date in milliseconds. |
| data.hostedUrl | string | Hosted MP4 link, returned only when the request set hostVideo. It plays without TikTok's signed CDN URL and without any cookie, and it stops working at expiresUtc. |
| data.id | string | TikTok video id. |
| data.language | string | Detected spoken language of the audio (BCP-47 style code, e.g. "en"). |
| data.ownerUsername | string | Creator handle, without the leading @. |
| data.segments | object[] | Timed transcript segments in playback order, one per sentence, so a segment locates a specific line in the video rather than a whole speaker turn. Populated whenever the provider has data for the entity. |
| data.segments[].endSeconds | number | Segment end offset in seconds, taken from the last word it contains. |
| data.segments[].language | string | Detected language for this segment. |
| data.segments[].speaker | string | Speaker label for this segment, stable within one response and meaningless across responses. Telling voices apart is a guess, not an identification, and the label is not a name. |
| data.segments[].startSeconds | number | Segment start offset in seconds, taken from the first word it contains. |
| data.segments[].text | string | Text of this segment. |
| data.segments[].words | object[] | Per-word timings for this segment, returned only when the request set wordTimestamps. Words carry no confidence score: the recognizer scores a phrase rather than a word. |
| data.segments[].words[].endSeconds | number | Word end offset in seconds. |
| data.segments[].words[].startSeconds | number | Word start offset in seconds. |
| data.segments[].words[].text | string | The recognized word, in display form with its own punctuation. |
| data.source | "audio_asr" | "transcript_unavailable" | How the text was produced. "audio_asr" means the words come from speech recognition over the video's audio, never from a caption track TikTok published - for TikTok's own captions, use tiktok.video_transcript. "transcript_unavailable" means recognition did not complete for this video, so the transcript is empty for that reason rather than because the video has no speech in it; the rest of the record is still what we resolved, and no audio time is charged. Populated whenever the provider has data for the entity. |
| data.thumbnailUrl | string | Cover image for the video. A signed, short-lived TikTok CDN URL, often served as HEIC rather than JPEG, so fetch it promptly and transcode if you need broad browser support. |
| data.transcript | string | Full spoken-word transcript, recognized from the video's audio track. Populated whenever the provider has data for the entity. |
| data.url | string | Canonical URL of the video this transcript came from. |
| found | boolean | |
| reason | "not_found" | Present only when `found` is false, and says why there is no result. `not_found`: the source states the target does not exist, or returned nothing for it. A `found: false` answer is a successful call, not an error, and `costUsd` is what it actually cost. |
| Preço | ||
| Preço por requisição (teto) | USD | US$ 0,095 |
| Preço /mil minutos de áudio | USD | US$ 5,94 |
Dúvidas frequentes
Sobre o endpoint Transcrição de vídeo (versão completa) da API para TikTok
O endpoint Transcrição de vídeo (versão completa) da AnyAPI para TikTok retorna dados de TikTok em JSON normalizado com uma chamada POST para /v1/run/tiktok.video_transcript_full. Transcreva o áudio falado de um vídeo do TikTok com o speech-to-text próprio da AnyAPI: segmentos com tempo por frase, rótulos de locutor, idioma detectado e tempos por palavra opcionais, para vídeos em que o TikTok não publica faixa de legenda e para vídeos cuja faixa de legenda erra as palavras. A AnyAPI baixa o vídeo e processa o áudio com o MAI-Transcribe-2 em vez de ler qualquer coisa que o TikTok escreveu, e é por isso que lida com fala em outros idiomas além do inglês e ouve palavras que as legendas nativas entendem errado. A resposta também traz o ID do vídeo, a URL, o nome de usuário do criador e a imagem de capa. Ative hostVideo para também receber o MP4 em um link hospedado que toca sem depender da URL de CDN assinada e de curta duração do TikTok, que expira. Se você só quer o que o próprio TikTok publicou, e por um décimo do preço, use tiktok.video_transcript. A AnyAPI retorna um schema normalizado, seja qual for a fonte que atende. Custa US$ 5,94 por mil minutos de áudio, e nunca mais de US$ 0,095 por requisição, em dólares, sem assinatura e sem mínimo mensal. Nos últimos 30 dias, 90,1% das chamadas ao endpoint Transcrição de vídeo (versão completa) da API para TikTok feitas pela AnyAPI tiveram sucesso, com tempo de resposta mediano de 4,7 segundos, em 172 chamadas medidas.