Turn your podcast or long video into ready-to-post Shorts
Turn recorded podcasts, interviews and talking videos into vertical Shorts. Find useful moments, review the cuts, follow faces near selected speaker positions and add animated captions.
ETA product guide · Updated
Workflow illustration · not generated footage
Find useful moments without reviewing every minute manually
Keep control of clip boundaries, captions and speaker framing
Reuse original footage and voices in downloadable vertical videos
How to create your video
ETA’s podcast-to-Shorts workflow repurposes your existing recording instead of generating a new presenter or replacing anyone’s voice. Upload a video podcast, interview, lesson or other talking video. The system transcribes the speech with word-level timing, then suggests a small set of useful moments. You decide what to keep. Review the original recording, adjust the beginning and end, correct captions and set the framing before a separate export stage creates your vertical videos. This is an assisted editing workflow for creators who want repeatable controls, not a promise that every suggested clip will perform well.
- Step 1
Upload a recording you have permission to use
Start with an MP4, MOV or WebM video containing clear spoken audio. This release accepts 30 seconds to 30 minutes of footage, up to 200 MB and 1080p. Compress larger recordings before uploading. Uploads go directly to private storage in chunks, with automatic retries for temporary network interruptions. Add a descriptive project name, the spoken language or auto detection, and optional context about the ideas you want to highlight. Confirm that you have permission to process the speakers’ appearances and voices.
- Step 2
Analyze the speech and review suggested highlights
Start analysis only when the recording is ready. Whisper on fal transcribes speech and returns word timestamps and speaker labels. Gemini through fal uses the transcript to propose up to five distinct, self-contained moments. Each suggestion includes a title and a reason for the selection. The target length is a preference, not a guarantee: fewer suitable clips may be returned. Analysis costs 40 credits and is separate from export. Completed analysis remains saved if you choose not to render immediately.
- Step 3
Choose the cuts, framing and caption treatment
Preview each cut in the original recording and adjust its start and end within the 15–60 second range. Keep questions, qualifications and conclusions intact so a short does not misrepresent the speaker. Correct caption words without changing the original audio. Choose animated word highlighting, clean captions or no burned-in captions. For framing, use assisted face-follow, a fixed manual crop or a full-frame layout with a blurred background. In multi-person recordings, listen to each voice and map its horizontal screen position before exporting.
- Step 4
Export selected clips and check the finished files
Select only the clips you want. Export costs 10 credits per selected clip, including a new render after edits. ETA cuts the original footage, applies your framing and captions, and creates a 720 × 1280 vertical MP4 with a separate SRT caption file. Jobs continue in the background and completed files remain attached to your private project. Watch every export for framing, subtitle accuracy and context before publishing. Download the files and upload them yourself to YouTube Shorts, TikTok, Instagram Reels or another platform.
What you can create
- Private resumable MP4, MOV and WebM uploads up to 200 MB, 1080p and 30 minutes
- Whisper transcription with word timestamps and speaker labels through fal
- Gemini on fal suggests up to five self-contained highlights
- Editable 15–60 second cuts, titles and caption-word corrections
- Assisted face-follow with voice-to-position mapping, manual crop or full-frame fit
- Animated word-highlight or clean captions, 720 × 1280 MP4 and separate SRT exports
- Saved projects, selected-clip rendering, background jobs and cancellation
Know the limits before you generate
- Requires a video recording with spoken audio, 30 seconds to 30 minutes long. Audio-only podcasts and link imports are not supported in this release.
- Face-follow uses detected faces near your chosen anchor. It is not guaranteed active-speaker recognition: match voices to positions for multi-person recordings and review camera changes.
- Transcripts and suggested moments need human review. Caption corrections do not alter the original spoken audio. No promised views, virality or search position.
- Analysis costs 40 credits; export costs 10 credits per selected clip. Completed analysis remains charged when you choose not to export. Re-rendering costs credits again.
- Live processing requires the Shorts migration, deployed worker with OpenCV/FFmpeg, fal access and account credits. No automatic social publishing or scheduling.
Ideas to get you started
A podcast answer that stands on its own
Find a guest’s useful answer, then keep just enough of the question to make it understandable outside the full episode. Use the source preview to avoid cutting an important qualification. For a two-person wide shot, map the voices to their left or right screen positions, or retain both people using full-frame fit. Review any point where the camera angle changes.
One practical lesson from a longer tutorial
Repurpose a short explanation or clearly stated tip from a lesson. Keep the full-frame layout when a slide, demonstration or on-screen diagram matters more than a close-up of the presenter. Captions help convey spoken instructions, but they do not replace important visual context. Include additional context in your own social post when a short cannot show the whole process.
A repeatable interview editing workflow
Save one project for each recording and use the suggested moments as a starting shortlist. Review cuts and captions, then export only the approved clips. Record which topics your audience actually responds to instead of treating an AI suggestion as a performance prediction. Original footage and voice remain the basis of each clip.
Questions, answered
What is a podcast-to-Shorts tool?
It is an editing workflow that turns parts of a longer recording into short vertical videos. ETA combines speech transcription, AI-assisted moment selection, editable cuts, framing controls and timed captions. It repurposes existing footage rather than creating synthetic speakers or rewriting the recorded conversation.
Does ETA automatically track whoever is speaking?
ETA offers assisted face-follow, not guaranteed active-speaker recognition. Audio speaker labels tell the system which voice is present; they do not identify a face. You can map each voice to its horizontal screen position. The renderer follows a detected face near that position and falls back to your anchor when detection is uncertain. Profiles, overlapping speech, occlusions and changing camera layouts need manual review.
Can I upload an audio-only podcast or paste a YouTube link?
Not in this release. Upload a video recording with spoken audio as MP4, MOV or WebM. The file must be 30 seconds to 30 minutes long, no larger than 200 MB and no higher than 1080p. If your episode is longer, prepare a shorter source segment first. Use recordings you own or are authorized to process and republish.
Can I change the captions and clip boundaries?
Yes. Review the original source, change start and end times, edit titles and correct individual caption words before rendering. Each clip must remain between 15 and 60 seconds and within the source duration. Caption edits affect displayed text only; they do not change what the speaker actually says. Save edits before export, and check the completed MP4 as the final reference.
What do I receive, and does ETA publish it for me?
You receive a vertical 720 × 1280 MP4 and a separate SRT caption file for each selected clip. Choose animated word-highlight captions, clean captions or no burned-in captions. ETA does not connect to or automatically post on your social accounts in this workflow. You review and upload the downloads yourself, subject to each platform’s requirements.
How are credits and cancelled jobs handled?
Analysis reserves 40 credits; rendering reserves 10 credits per selected clip. Successful stages remain charged. A failed or cancelled job returns that job’s reservation after the worker stops, while previous completed work is preserved. Repeated render requests for an already-running stage reconnect to the saved job rather than reserving again. Provider work already in progress may not stop instantly.
Will AI-selected clips get more views?
No tool can guarantee views or virality. ETA suggests moments based on the supplied transcript and context, not verified audience performance. Use the suggestions to reduce searching time, then apply your editorial judgment. Strong source material, a clear opening, accurate captions, sensible framing and a useful takeaway matter more than a predicted score.