Founder:pay once, never think about it again. Every feature we ship from here on is included.33 seats leftClaim a founder seat

Podcast audiogram maker

Makes a shareable video out of audio, with artwork and captions.

The short answer

An audiogram is a video made from audio: artwork or a still on screen, some motion driven by the sound, and captions carrying the words. It exists because audio cannot be posted to a feed and be watched. Vid2Shorts builds one from a clip of your episode, with fourteen visualizer presets and captions timed word by word, and every preset works on the free plan.

First Short free. No account, no software, nothing to upload.

This is the working tool, not a description of one. The first Short is free and needs no account. Everything on this page describes what the product does today, checked August 2026.

Why audiograms exist

A feed will not play your audio

Podcast audiences do not grow inside podcast apps. They grow where people already are, which is a scrolling video feed, and a feed has nowhere to put an MP3. An audiogram is the workaround: it wraps the audio in something watchable so the clip can live in the same places everything else lives.

The important consequence is that an audiogram is a video and has to be judged as one. It competes with clips that have faces and cuts and movement. Static artwork with a waveform bouncing along the bottom was novel some years ago and now reads as a placeholder, which is why the motion and the captions matter more than the fact that the format exists.

The captions carry it

The words are the picture

On an audiogram there is no performance to watch, so the captions are not an accessibility addition, they are the main visual element. That changes how they should be treated. They want to be large, they want to be timed per word so the eye is following something rather than reading ahead, and they want emphasis on the words that land.

Word by word timing does more work here than anywhere else, because it is the only thing on screen with a pulse. A whole line appearing at once on an audiogram is close to a still image with text on it, and the difference in how long somebody watches is not subtle.

Getting the words right therefore matters more too. Retype any line the transcriber misheard, or if you are working from a script or show notes, paste the real words and have them timed to the audio rather than transcribed. A misspelled guest name is a worse look on an audiogram than anywhere else, because it is the only thing to look at.

The motion

Movement that comes from the audio, not from a loop

The motion here is driven by the sound rather than being a decorative loop playing underneath it. Artwork pulses with the low end, spectrum rings and bars move with the frequency content, and the result tracks the actual audio rather than merely happening near it. Fourteen presets cover the range from nearly still to fully kinetic.

Which one to use is mostly a question of what the clip is. A thoughtful answer to a hard question wants something close to still, because busy motion under a serious point reads as unserious. A laugh, an argument, a moment of energy can carry a much more kinetic preset. The presets are all available on the free plan, so trying three costs nothing but the two minutes.

This is also the part of the toolkit that no comparable product has. Checking every tool people shop against us, none of them, including the largest and most capable editor on the list, has a music or audio visualizer at all.

Finding the clip

Which ninety seconds of a ninety minute episode

The hardest part of podcast promotion is not making the audiogram, it is deciding what goes in it. If the episode exists as video too, paste it and the whole transcript is read and ranked, with each candidate scored on whether it hooks, whether it carries emotion, whether any line is quotable and whether it stands alone without setup.

That last score is the one that matters most for podcasts specifically, because conversation is full of moments that were brilliant in context and mean nothing out of it. A tool that shows you why it picked something lets you catch that in seconds rather than after you have already made the clip.

Step by step

How to do it

01

Bring in the episode

A link or an upload. If the episode exists as video, the whole transcript gets read and the moments that stand alone come back ranked.

02

Choose the moment

Judge it on whether it works with no setup, which is where most podcast clips quietly fail.

03

Pick a visualizer preset

Fourteen of them, from nearly still to fully kinetic, all available on the free plan. Match the energy of the clip rather than the energy of the preset.

04

Caption it properly

Word by word and large, because on an audiogram the words are the picture. Fix any misheard name before it goes out.

05

Export vertical

1080 by 1920 MP4, ready for Reels, Shorts or TikTok, plus an SRT or VTT if you want a subtitle file too.

The hard facts

Specifications

The numbers rather than the adjectives, including the limits. If one of these is a dealbreaker it is better learned here than after paying.

Visualizer presets14, all included on the free plan
Motion sourceDriven by the audio, not a decorative loop
CaptionsWord by word, with emphasis on the strongest words
Output1080 by 1920 MP4, plus optional SRT or VTT
Runs inA browser, including on a phone
Output1080 by 1920 MP4, H.264, on every tier including free
Runs inAny modern browser. Nothing to install, works on a phone
Free tierOne Short a day, up to 30 seconds, with a small corner watermark
Account neededNot for your first Short. Captions and subtitle files need a free account
Questions

What people ask about this

What is a podcast audiogram?
A video made from audio so it can be posted somewhere that will not play an MP3. Usually artwork or a still on screen, some motion driven by the sound, and captions carrying the words. It exists entirely because podcast audiences grow in video feeds rather than in podcast apps.
How do I make an audiogram for free?
Bring in a clip, choose one of the fourteen visualizer presets, caption it, and export. The free tier gives one Short a day up to thirty seconds with a small corner watermark, and every preset is included rather than held back for a paid plan.
Do audiograms actually work?
They work when they are treated as video and not as a compliance exercise. A static image with a waveform bouncing under it reads as a placeholder now. Large captions timed per word, motion that genuinely follows the audio, and a clip chosen because it stands alone are what separate one that gets watched from one that gets scrolled.
What should be on screen in an audiogram?
Something that identifies the show, and the words. Cover art or a guest photo is enough for the first, and the captions do the second and carry most of the attention. Resist filling the frame with logos and handles, because every element competing with the words costs you the words.
How long should a podcast clip be?
Short enough that nothing in it is waiting, which usually means shorter than the moment felt when you heard it. The test worth applying is whether the clip needs setup that is not inside it. A brilliant answer to a question nobody heard is half a conversation, not a clip.
Can I use my own artwork?
Yes, and it is the natural thing to put on screen. The motion is driven by the audio rather than by a preset loop, so the artwork pulses with the low end rather than merely sitting there while something else animates near it.
Do other tools make audiograms?
Very few of the tools people compare us against have a visualizer at all. Checking every one of them, including the largest and most capable editor in that comparison, none of them offers this. If audio is a real part of what you publish, it is the clearest single reason to be here.

Try it on one video

The first Short is free, needs no account, and takes about two minutes. That answers more than any page about a tool can, including this one.

Open the Studio

Keep going