The footage already knows who said what.
Every person was recorded on their own track and every word is transcribed against it, so the edit is deleting text and the clips are lines you highlight. No fourth subscription to move the files into.
Free during the beta. Your files stay downloadable, with no expiry.
A finished take listed as separate files: the composed program plus one 1080p master per speaker.
Transcribed per speaker, because the recording was per speaker
Most transcription starts from a flattened mixdown and then guesses at who was talking, which is why two people agreeing at the same time comes back as one confused paragraph. Here there is nothing to guess: each participant had their own local recording, so the transcript is built against known tracks with the right name already on every line.
You get a speaker-labelled transcript for the session and caption files that match it, timed against the video rather than approximately aligned to it.
- One speaker per track. No diarisation guesswork, no “Speaker 2” to rename afterwards
- Clean crosstalk. Overlapping speech stays attributed to the person who said it
- Captions that match. Subtitle files generated from the same transcript the edit uses
- Searchable. Find the moment by the sentence, not by scrubbing
The transcript isn’t a document about the video. It is the video, written down.
A transcript with each line labelled by speaker name, one line highlighted and one deleted.
The edit is a text document you delete from
Nothing here needs a timeline, a scrub bar or an afternoon of tutorials. It needs you to read the conversation back and take out what you don’t want.
Edit by deleting text
Delete a sentence and the video loses it. Cut the “um”s in a pass, or take out a whole section that didn’t land. It is the same gesture either way.
Splice and reorder takes
A session can be many takes. Assemble them, put the second answer before the first, drop the one where the doorbell went, all without opening a timeline editor.
Show notes and chapters
Notes, chapters and title options drafted from what was actually said, so the description writes itself from the conversation instead of your memory of it.
Caption files
Subtitles exported alongside the video, generated from the speaker-labelled transcript rather than re-derived from the audio.
Audio-only mixdown
The audio version of the episode as a single file, for the podcast feed that doesn’t want your camera.
Publish to YouTube
Send the finished video straight out, with the title, description and chapters you just drafted attached to it.
The good bits, found by reading rather than watching
A ninety-minute conversation has maybe six moments worth posting, and finding them by scrubbing is the reason most of them never get posted. The transcript is a better search surface than a waveform: highlights are picked out of what was said, and you keep or discard each one.
A clip comes out cropped to the speaker, with captions burned in and your branding on it, in whichever shape the place you’re posting expects.
- Picked from the transcript. Suggested moments you can accept, trim or overrule
- Cropped vertical. Framed on whoever is talking, not a letterboxed 16:9 in the middle of a phone screen
- Captions burned in. Because the feed plays it on mute first
- 9:16, 1:1 and 16:9 exports. One clip, every shape you need to post it in
Clips are the part everyone intends to do and nobody has an hour for. Reading beats scrubbing.
Three vertical 9:16 clip cards cut from one session, each cropped to the speaker with burned-in captions.
What one session can hand you
All of it from the same recording, in the same place it was recorded. No export, no upload, no second login.
Video
- The composed program as a plain MP4
- Every participant’s own 1080p track, separately
- The edited cut, assembled from the transcript
- Vertical, square and wide clip exports
Words
- A speaker-labelled transcript for the session
- Caption files that match the finished cut
- Show notes and chapter markers
- Title options drafted from the conversation
Audio and delivery
- An audio-only mixdown for the podcast feed
- Per-track files for anyone mixing properly
- Export packs, downloadable any time
- Publish straight to YouTube
And if you’d rather cut it somewhere else, take everything with you
None of this is a lock. The separate tracks are plain files you can download whenever you like and open in whatever you already use, and the transcript and caption files come with them. Editing here is meant to be the fastest route, not the only exit.
What we won’t do is put your footage behind a countdown. Recordings don’t get deleted because a fortnight passed, and downloading a separate track isn’t an hour off a meter.
- Plain MP4s. No export queue, no processing tier, no expiry date
- Per-track downloads that aren’t metered by the hour
- Transcript and captions exported as files, not just displayed in an app
- Your recordings outlive your subscription. No delete window counting down
A finished take listed with the composed program file and one downloadable file per participant, alongside transcript and caption exports.
Recording, streaming, editing and clips are four separate subscriptions everywhere else, around $80 to $100 a month across four logins.
They’re four things one room already has the material for. The tracks are here, the transcript is here, the branding is here. Charging you again to move the same footage between them would be a business model, not a feature.
About the edit
Do I have to edit here?
No. Every take downloads as plain MP4s: the composed program, and each participant’s own 1080p track. Cut them in whatever you already use. The transcript and caption files export too. This is the fast path, not a walled one.
How does deleting text delete video?
Because the transcript is timed against the recorded tracks, every word maps to a moment. Remove the words and the cut follows them. It’s the same edit a timeline would make. You just make it by reading instead of scrubbing.
Are the transcripts actually labelled by person?
Yes, and not by guessing. Each participant records on their own track, so a line is attributed by which track it came from. That’s also why people talking over each other stays legible instead of collapsing into one speaker.
Can I fix a clip the AI picked badly?
Suggestions are a starting point, not a verdict. Trim the in and out points, change the crop, re-word the caption, or throw the clip away and highlight the passage you actually wanted.
What formats do clips export in?
9:16 for vertical feeds, 1:1 for square, and 16:9 when it’s going back on a normal video page. Same clip, framed for each, with captions burned in.
Does editing count against some kind of credit?
No. There are no AI credits, no separate editor seat and no per-hour charge for a transcript or a clip. One flat price, and pricing published before anyone is asked to pay.
Finish the next one the same day you record it
We’re letting people in a few at a time, and the beta is free. Leave your email and we’ll open a studio for you.