VidHelm is a free, open-source desktop video editor built to be flown with an AI co-captain. Cut the dead air, narrate in one take, drop sound effects on your beats, spin your 3D prints over your footage, and export YouTube-ready, by hand, or by asking.

Import → tag the beats → tighten → narrate → sound → export. Every step has a fast path.
Finds every silent gap, or motionless stretch in silent footage, and splices it out with tiny crossfades. A rambling take becomes a tight one in one click.
Tap M at every beat, the joke, the reveal, the pop. Tags are the language of the whole app: clips snap to them, SFX land on them, narration lines sync to them.
Paste your script and hit record: the video plays while your lines light up in time. One continuous take, no clicking between lines, lands straight on the voice track.
A setup wizard records about twenty seconds of your voice and installs a free local engine for you: XTTS-v2, or audio.cpp if you want no Python and licence-clean models. Lines land on your tag points automatically.
Thirteen classic effects synthesized on your machine, your own sounds dropped into a folder, or press record and make one yourself. Describe a sound in words and a local model can generate it too.
Drop in an STL, 3MF, OBJ or GLB, pose it, pick a colour and finish, and render a spinning turntable straight onto the timeline. Pick the green screen backdrop and your print floats over your footage.
Point VidHelm at one folder and every sub-folder inside it becomes a project. Dropping files in with Explorer is the import. Saving writes back to the same folder, so a project is something you can copy or hand to someone.
On-device Whisper transcribes the whole timeline into styled captions, phrase or word-by-word karaoke mode. Nothing leaves your computer.
Labeled voice / SFX lanes, drawable volume automation, and one checkbox that masters the mix to YouTube's −14 LUFS target with peaks in check.
Every export is auto-checked, loudness, true peak, black frames, sampled stills, so you never upload a broken render again.
Sample frames from your video, pick one, type a catchy subtitle, out comes a 1280×720 thumbnail with your logo composited on.
Video, audio and image formats are decided by FFmpeg rather than a fixed list, so unusual files still work. Anything unusable is refused with a plain-English reason instead of landing as a broken clip.


VidHelm ships an MCP server. Point Claude (or any agent) at it and your AI gets 25 tools to drive the running app, while you watch it happen in your window.
Your standing instructions, run on every new video. Toggle steps on and off (off = #-commented, just like G-code), free-type anything, your AI reads it and does the busywork.
The slowest part of editing a talking-head video is hunting down every "um", breath and dead-air gap. VidHelm does that pass in one click, or you can just ask your AI to do it.
VidHelm scans the whole timeline for silent gaps and, in silent footage, motionless stretches, then splices them out with short crossfades so the cuts stay seamless. A twenty-minute rambling take turns into a tight edit before you touch anything else.
No subscription, no watermark, no render credits. VidHelm is MIT-licensed and open source, ships as a Windows installer with FFmpeg bundled, and every AI feature, captions, voice cloning, sound generation, runs locally on your own machine.
Connect Claude, Cursor, VS Code, Ollama or LM Studio to the built-in MCP server and say "cut the pauses, caption it and export for YouTube". Twenty-five tools drive the real app while you watch it happen in your window.
Auto captions from on-device Whisper, cloned-voice narration on your tag points, a 1280×720 thumbnail builder, and exports mastered to YouTube's −14 LUFS target with a quality report on every render.
Windows installer with FFmpeg bundled, nothing else to install. Or run it from source anywhere Node runs.
Developers:
git clone https://github.com/RandoTechNerd/VidHelm cd VidHelm && npm install && npm run dev