How it works

How an AI radio DJ works: scripts, voices and ducking

By Nir Kalfon · Updated · 7 min read

An AI radio DJ works in five steps: it figures out which song is about to play, finds a few reliable facts about it, has a language model write a short intro in the DJ's style, turns that text into speech, and lowers the music while the voice plays. The hard part isn't any single step. It's doing all five in the few seconds between two songs, without making things up.

Here's how each step works in Stationz.fm, a Chrome extension that adds an AI DJ to YouTube playlists, and how it compares with the best known AI DJ, Spotify's.

Step 1: Figure out the song

On Spotify, the app knows exactly which track is playing. On YouTube it's messier: a video is just a title and a channel name, like "Artist - Song (Official Video) [Remastered]" uploaded by some channel.

So the first job is to turn that into an artist and a song title:

  • Rules first. Most music video titles follow common patterns, and simple rules can split them into artist and title, dropping the "Official Video", "HD" and "Lyrics" noise.
  • A language model when the rules aren't sure. For messy titles, an AI model reads the title and channel and returns its best guess, along with how confident it is and whether the video is music at all.
  • Remember the answer. Once a video is resolved, the answer is saved, so the next listener doesn't need the same lookup.

Two safety rules come out of this step. If the video isn't music (a podcast, a vlog), the DJ stays quiet. And if the confidence is low, the DJ is told not to state any facts at all: it only introduces the video by its title. A short, plain intro is better than a confident wrong one.

Step 2: Pick facts that are checked

Language models are good at sounding sure about things that never happened. A music DJ that invents stories is worse than no DJ.

So Stationz doesn't rely on what the model "knows". It keeps a music facts library built from Wikipedia and MusicBrainz. Short facts about songs, albums, artists and the people behind them are pulled from those pages, and the page match is checked first, because the wrong Wikipedia page means wrong facts. When a song isn't in the library yet, it's queued to be researched, and meanwhile the DJ can use checked facts about the artist if the artist is already there. On top of that, there are long researched stories for many well-known artists (see the story behind every song).

When it's time to introduce a song, the DJ picks from that library:

  • How many facts depends on the talk level. Minimal uses none (just the song and artist). Balanced uses one short fact. Talkative uses up to two longer ones.
  • Facts you've heard are skipped. The library remembers which facts each listener has already heard, and picks fresh ones while there are any left for that song.
  • Sad facts are handled gently. When a fact touches on death, illness or tragedy, the DJ is told to keep a calm, respectful voice with no jokes or hype.

Step 3: Write a short script

Now a large language model writes the actual intro. It gets a set of instructions, roughly:

  • Who it is. The DJ's name, personality and style. Each of the DJs has its own.
  • How long. Minimal is one short sentence of a few seconds. Balanced is two or three sentences, up to about 10 seconds. Talkative is a few sentences, up to about 20 seconds.
  • The tone. Calm, hype, nerdy, poetic, academic and others.
  • The facts it may use, and only those. The instruction is explicit: work these facts in, in your own words, and don't add any other names, dates or numbers.
  • The language. If you listen in Hebrew, Japanese or any of the 16 languages (languages other than English are on Pro), the DJ is told to talk like a native radio host, not to translate. More in An AI DJ that speaks your language.
  • Sound spoken, not written. Short and long lines mixed, no stock DJ phrases, and punctuation chosen so the voice doesn't pause awkwardly.

Real radio hosts don't say the same thing every time a song comes around, so neither should the DJ. The last few intros for a song are kept, and the new one is compared with them word by word. If it's too similar, the model is asked again for a completely different angle.

There are also small radio touches: now and then the DJ reacts to the song that just ended, and on Pro it can give the local weather or mention a notable date between songs.

Step 4: Turn text into a voice

The script becomes speech in one of two ways:

Free DJs (Emma, Leo)Pro DJs (10 voices)
Voice engineChrome's built-in text to speechGoogle's Gemini AI voices
Where it runsOn your computerGenerated in the cloud, then played in the tab
How it soundsClear, but recognizably syntheticExpressive, like a real radio host
LimitUnlimited120 minutes a month on Pro, 400 on Pro Max

For the free DJs, the extension picks the best voice Chrome offers on your computer for your language and the DJ's gender, preferring the higher quality "natural" voices when they're installed. For Pro, each DJ has its own Gemini voice and a style description (for example "calm nighttime storytelling" or "confident British storyteller"), and in other languages the voice is asked to speak with a native accent.

If a Pro voice can't be generated in time (the service is busy, say), the DJ falls back to the free voice rather than skipping the intro.

Step 5: Duck the music and time the intro

On real radio, the host talks over the start of the song while the music sits lower. That's called ducking, and it's what makes an AI DJ feel like radio instead of a voice note.

Stationz does it inside the YouTube tab:

  • The volume fades down, not off. When the DJ starts, the song's volume fades down over about a third of a second, by default to about a quarter of your volume. When the DJ finishes, it fades back up over about half a second. Fades are done in small steps so you hear a smooth dip, not a click.
  • Your volume wins. If you change the volume yourself while the DJ is talking, the extension notices and doesn't overwrite it.
  • Never during ads. If a YouTube ad is showing, the music isn't ducked.

Timing is the other half:

  • Intros are prepared ahead. While one song plays, the intros for the next two songs are already being written and voiced, so most are ready the moment the song changes.
  • A new song starts quiet. When the next video starts loading, it starts at the lowered volume, so you don't get a few seconds of full blast and then a sudden dip for the DJ.
  • A short wait if needed. If an intro isn't ready, the song waits at its start for up to about three seconds. The very first song of a station waits for the welcome, since that's when the DJ introduces itself.

How Spotify DJ does it

Spotify's DJ, launched in February 2023, is built from similar parts, made in-house. According to Spotify's launch announcement, it combines Spotify's personalization, a voice made with dynamic voice technology from Sonantic (a company Spotify acquired), and commentary written with OpenAI's generative AI together with Spotify's music editors (Spotify Newsroom, as of October 2026). The English voice is modeled on Spotify's Xavier "X" Jernigan (Wikipedia).

Where Spotify is stronger: it's built into Spotify's own app and catalog, it knows exactly what's playing, and it chooses a personal lineup for you from your listening history. Where the approaches differ: Spotify DJ picks the music itself, while Stationz narrates over a YouTube playlist you choose, and Stationz lets you pick the voice, the talk level and the tone. For a full side by side, see Stationz vs Spotify DJ.

FAQ

Does the AI DJ make up facts?

It's designed not to. When the library has checked facts for a song (from Wikipedia and MusicBrainz), the DJ is given those and told not to add any others. When it isn't confident which song is playing, it's told not to state any facts at all.

Why does the free DJ sound different from the Pro DJs?

The free DJs use Chrome's built-in voices on your computer. Pro DJs use Gemini AI voices, which are more expressive and sound closer to a real radio host.

Does the music stop while the DJ talks?

No. It dips to a lower volume and comes back up when the DJ is done, like on the radio.

Is the same intro played every time?

No. Each intro is written fresh, and it's checked against the last few intros for the same song so the DJ says something different.

Hear it on your own playlists

Stationz adds an AI radio DJ to YouTube in desktop Chrome. Free to start, 16 languages with Pro.