Podcast audiogram maker
Makes a shareable video out of audio, with artwork and captions.
An audiogram is a video made from audio: artwork or a still on screen, some motion driven by the sound, and captions carrying the words. It exists because audio cannot be posted to a feed and be watched. Vid2Shorts builds one from a clip of your episode, with fourteen visualizer presets and captions timed word by word, and every preset works on the free plan.
First Short free. No account, no software, nothing to upload.
This is the working tool, not a description of one. The first Short is free and needs no account. Everything on this page describes what the product does today, checked August 2026.
A feed will not play your audio
Podcast audiences do not grow inside podcast apps. They grow where people already are, which is a scrolling video feed, and a feed has nowhere to put an MP3. An audiogram is the workaround: it wraps the audio in something watchable so the clip can live in the same places everything else lives.
The important consequence is that an audiogram is a video and has to be judged as one. It competes with clips that have faces and cuts and movement. Static artwork with a waveform bouncing along the bottom was novel some years ago and now reads as a placeholder, which is why the motion and the captions matter more than the fact that the format exists.
The words are the picture
On an audiogram there is no performance to watch, so the captions are not an accessibility addition, they are the main visual element. That changes how they should be treated. They want to be large, they want to be timed per word so the eye is following something rather than reading ahead, and they want emphasis on the words that land.
Word by word timing does more work here than anywhere else, because it is the only thing on screen with a pulse. A whole line appearing at once on an audiogram is close to a still image with text on it, and the difference in how long somebody watches is not subtle.
Getting the words right therefore matters more too. Retype any line the transcriber misheard, or if you are working from a script or show notes, paste the real words and have them timed to the audio rather than transcribed. A misspelled guest name is a worse look on an audiogram than anywhere else, because it is the only thing to look at.
Movement that comes from the audio, not from a loop
The motion here is driven by the sound rather than being a decorative loop playing underneath it. Artwork pulses with the low end, spectrum rings and bars move with the frequency content, and the result tracks the actual audio rather than merely happening near it. Fourteen presets cover the range from nearly still to fully kinetic.
Which one to use is mostly a question of what the clip is. A thoughtful answer to a hard question wants something close to still, because busy motion under a serious point reads as unserious. A laugh, an argument, a moment of energy can carry a much more kinetic preset. The presets are all available on the free plan, so trying three costs nothing but the two minutes.
This is also the part of the toolkit that no comparable product has. Checking every tool people shop against us, none of them, including the largest and most capable editor on the list, has a music or audio visualizer at all.
Which ninety seconds of a ninety minute episode
The hardest part of podcast promotion is not making the audiogram, it is deciding what goes in it. If the episode exists as video too, paste it and the whole transcript is read and ranked, with each candidate scored on whether it hooks, whether it carries emotion, whether any line is quotable and whether it stands alone without setup.
That last score is the one that matters most for podcasts specifically, because conversation is full of moments that were brilliant in context and mean nothing out of it. A tool that shows you why it picked something lets you catch that in seconds rather than after you have already made the clip.
How to do it
Bring in the episode
A link or an upload. If the episode exists as video, the whole transcript gets read and the moments that stand alone come back ranked.
Choose the moment
Judge it on whether it works with no setup, which is where most podcast clips quietly fail.
Pick a visualizer preset
Fourteen of them, from nearly still to fully kinetic, all available on the free plan. Match the energy of the clip rather than the energy of the preset.
Caption it properly
Word by word and large, because on an audiogram the words are the picture. Fix any misheard name before it goes out.
Export vertical
1080 by 1920 MP4, ready for Reels, Shorts or TikTok, plus an SRT or VTT if you want a subtitle file too.
Specifications
The numbers rather than the adjectives, including the limits. If one of these is a dealbreaker it is better learned here than after paying.
| Visualizer presets | 14, all included on the free plan |
|---|---|
| Motion source | Driven by the audio, not a decorative loop |
| Captions | Word by word, with emphasis on the strongest words |
| Output | 1080 by 1920 MP4, plus optional SRT or VTT |
| Runs in | A browser, including on a phone |
| Output | 1080 by 1920 MP4, H.264, on every tier including free |
| Runs in | Any modern browser. Nothing to install, works on a phone |
| Free tier | One Short a day, up to 30 seconds, with a small corner watermark |
| Account needed | Not for your first Short. Captions and subtitle files need a free account |
What people ask about this
What is a podcast audiogram?
How do I make an audiogram for free?
Do audiograms actually work?
What should be on screen in an audiogram?
How long should a podcast clip be?
Can I use my own artwork?
Do other tools make audiograms?
Try it on one video
The first Short is free, needs no account, and takes about two minutes. That answers more than any page about a tool can, including this one.
Open the Studio
