Tools
Video to Markdown API
Submit any video file. Get back a readable Markdown document with the full transcript, duration, dimensions, and metadata.
Try it below:
Turn video files into Markdown
Upload a video file or submit a URL, get Markdown back.
What is Exabase's video to Markdown API?
Submit a video file to the Exabase Extract API and get a Markdown document back instead of JSON. The same extraction pipeline runs: audio track extraction, transcription, metadata parsing, and thumbnail generation. The difference is the output. Pass ?format=markdown when retrieving the result and you get a human-readable Markdown document with the full transcript instead of a JSON object.
The Markdown output includes video duration, frame dimensions, MIME type, file size, and the complete transcript as continuous text. MP4, MOV, AVI, WebM, MKV, and other major formats are supported. No FFmpeg, no speech-to-text model, no GPU.
What you get back
The Markdown response is a single text document. Duration and frame dimensions appear in the metadata block. The full transcript follows as continuous prose. The result is ready to feed directly into an LLM context window, store as a readable summary, or use as source material for content repurposing.
You can also configure webhooks with webhookFormat: "markdown" so completed jobs POST the Markdown document directly to your server.
One multi-modal API
The same POST /v2/extract endpoint handles every content type Exabase supports, and the ?format=markdown parameter works across all of them. For video, you get duration, dimensions, and the full transcript. For audio, you get duration and the transcript. For PDFs, you get document metadata and text. For images, you get dimensions and OCR text. For web pages, you get the title, site name, and content.
The submission flow is identical across all types: one endpoint, one SDK method, one webhook configuration. The Extract docs cover each content type in detail.
What you can build with it
Feed video transcripts directly into your agent's context window as readable text. Build a video search platform where transcripts are stored as readable Markdown. Power a learning copilot that transcribes lecture recordings and makes them searchable.
Store transcribed video as Resources in a Base and the transcript becomes searchable through Deep Search. Workers can process new recordings as they arrive.
Beyond Markdown
The Extract API also returns JSON (the default) with timestamped chunks for seek-to-source UIs and subtitle generation. Use Markdown for LLM context and human review, JSON for structured processing with timestamps. The Video to JSON tool page covers the JSON output.
How do I use it?
Get your API key
Free, no credit card:
Sign up at exabase.io and copy your API key from the dashboard.
Submit a video file
Using the SDK (Node.js):
Get Markdown back
Poll the job until state reaches completed, then request the Markdown format.
One API call. No FFmpeg, no speech-to-text model, no Markdown conversion.
Why Exabase
Works with other Exabase features
FAQs
What is Exabase?
Exabase is infrastructure for AI agents. It gives your agents memory, versioned file storage, AI deep search, and context automation through a set of APIs. Store what your agent learns, search inside any content type, and keep knowledge bases current automatically. Built for production use. Give your agent precise context and cut your token spend by up to 81%.
Who uses Exabase?
Developers and teams building AI agents, copilots, and RAG applications. If your agent needs to remember things between sessions, store and retrieve files, search across documents and media, or stay up to date without manual maintenance, Exabase handles that infrastructure so you can focus on your product.
How fast is processing?
Processing time depends on the file length, resolution, and format. Processing is asynchronous, so your application is not blocked while extraction runs. Poll for results or configure a webhook to be notified on completion.
How do I get Markdown instead of JSON?
Add ?format=markdown to the GET /v2/extract/{jobId} request. For webhooks, set webhookFormat: "markdown".
Does the Markdown output include timestamps?
The Markdown includes the full transcript as continuous text and the total duration. For per-segment timestamps, use the JSON format and fetch chunks.
Can I still get JSON?
Yes. JSON is the default. Both formats are available from the same extraction job.
How is this different from Video to JSON?
Same pipeline, different output. Video to JSON returns structured JSON with timestamped chunks. Video to Markdown returns a readable document.
What video formats are supported?
MP4, MOV, AVI, WebM, MKV, and other major formats.
Does it transcribe on-screen text?
No. The API transcribes the audio track. On-screen text is not extracted separately.
Do I have to poll for results?
No. Configure a webhook with webhookFormat: "markdown" to receive the document on completion.
How long are files retained?
Stored files are retained for 1 day from job creation, then permanently deleted. Download or copy anything you need before the retention window expires.
Is there an SDK?
Yes. The @exabase/sdk package for Node.js/TypeScript handles file streaming, job polling, and chunk retrieval. Install with npm install @exabase/sdk. Or call the REST API directly from any language.
What other content types support Markdown output?
All of them. See the PDF to Markdown, Audio to Markdown, Image to Markdown, and Website to Markdown tool pages.