← Lab notebook

My 17-track AI soundtrack failed—what I’ll change

I generated 17 instrumental beds locally with MiniMax Music 3 in ComfyUI on an Apple M1 Max, and by my own rough count I spent more than 25 hours on the batch.

I own and build this site and the AI Maker Lab channel — this is a build log, not an independent review.

My conclusion is that this 17-track batch did not pass my own taste test, even though it used three-part Structured Captions, so my second attempt will use 30-second, one-variable tests before a long render.

What I built: 17 local instrumental beds

The batch contains 17 original instrumental beds: five sized for Instagram Reels or Stories and twelve for YouTube episodes. I ran MiniMax Music 3 through ComfyUI on an Apple M1 Max with 64 GB of unified memory. This machine used the fp16 DiT weights, the pruned bf16 text encoder, and the DAV VAE. Every track also had a three-part Structured Caption, an instrumental Vocal Details block, and a timed Arrangement.

The library records a requested style, tempo, key, delivered duration, and intended use for each track. Those fields describe the request, not a verified musical analysis of the output. MiniMax’s model card explains that tempo, key, instrumentation, and structure are generative controls rather than guarantees, so listening remains the check. The WAV masters stay at 44.1 kHz and 16-bit, while the players below use measured MP3 publication copies with matching durations.

My own judgment is that none of these 17 tracks is good enough to keep for the channel. That verdict applies only to this batch and my taste. It is not a comparison with another model or a general judgment about MiniMax Music 3. The library’s own operating rule is also to audition every bed under actual narration rather than only by itself.

Listen for yourself

All 17 delivered beds are available here as MP3 copies. Each copy was encoded from its recorded delivery WAV, and its measured duration matched the ledger exactly. The style, tempo, key, duration, and use below come from that first-party ledger; tempo and key remain requested targets that require listening.

The MiniMax-Music3 licence requires public content to disclose clearly and prominently that it is machine-generated. The following standing line is this repository’s chosen disclosure wording, not wording prescribed by the licence: Music generated with MiniMax-Music3.

Five Instagram beds

ig-rock-launch — pop rock / alternative rock, 150 BPM, B♭ major · 38 s · for feature-launch reels
ig-rock-pulse — electronic-leaning pop rock, 111 BPM, C minor · 38 s · for problem-to-fix montages
ig-jazzhop-groove — Rhodes jazzhop, 86 BPM, A minor · 38 s · for maker-process reels and code timelapses
ig-boombap-head-nod — jazzy boom bap, 90 BPM, G minor · 38 s · for confident demo walkthroughs
ig-jazz-swing-cafe — piano-trio light swing, 118 BPM, B♭ major · 36 s · for office-BTS and wrap reels

Twelve YouTube beds

yt-rock-opener — post-grunge editorial guitar bed, 88 BPM, F♯ major · 1:04 · for episode openers and chapter transitions
yt-rock-momentum — indie rock bed, 91 BPM, F major · 2:12 · for build sequences
yt-rock-anthem-outro — bright anthemic pop rock bed, 143 BPM, A major · 2:17 · for results segments and outro CTA
yt-lofi-study — study-beat chillhop, 88 BPM, G minor · 1:46 · for primary long talking and coding bed
yt-jazzhop-sax-rain — sax-loop jazzhop, 86 BPM, E♭ major · 2:30 · for reflective essay segments
yt-jazzhop-rhodes — Rhodes jazzhop, 86 BPM, A minor · 1:28 · for long-form commentary
yt-boombap-city-night — jazzy boom bap, 90 BPM, G minor · 1:43 · for editing and montage segments
yt-jazz-bossa — bossa nova / smooth jazz, 133 BPM, D minor · 1:56 · for cafe and calm explanation segments
yt-jazz-cool-trio — cool jazz combo bed, 125 BPM, B♭ minor · 1:56 · for thoughtful walkthrough segments
yt-jazz-lounge-warm — lounge piano bed, 143 BPM, C major · 1:46 · for outro and recap segments
yt-8bit-arcade-level — chiptune / 8bit arcade level theme, 118 BPM, C major · 1:15 · for gameplay montages and build sequences
yt-8bit-arcade-night — chiptune / 8bit arcade night groove, 96 BPM, A minor · 55 s · for focused coding and deep-dive segments

Where my reported 25+ hours went

By my own rough count, I spent more than 25 hours on this batch. That is a self-report, not a total reconstructed from machine logs. The only wall-clock measurement here covers two 30-second drafts on this M1 Max. They took 40 minutes 46 seconds and 38 minutes 44 seconds, or about 80 seconds of wall time per second of audio on this machine.

That measured cost makes the order of tests important without proving any productivity outcome. The local operating recipe starts with a 30-second draft. Keep the seed fixed while changing one caption axis, such as tempo, palette, or production profile. Then freeze the caption, roll two or three seeds, and compare performances by listening. Commit one selected seed to the long render, setting duration to the intended length plus about ten seconds because generation may end early.

Six-node music iteration loop from a 30-second draft through caption and take tests to quality control and delivery
Two measured drafts on one M1 Max explain why the long render belongs after the caption and seed decisions.

The loop includes a real rejection path. Quality control checks the true peak and the loop seam before delivery, then sends a rejected bed back to the draft. That taste gate is separate from file completion. A reproducible render can still fail my own decision about whether the music belongs on the channel.

The researched workflow I should have used

The researched workflow is one chain: write a disciplined caption, test one variable, listen to soft targets, design the seam, and check the delivered level. A correct caption structure does not guarantee that I will like the result. The official announcement describes Structured Captions as temporal descriptions of emotion, instrumentation, vocals, and arrangement. My batch shows only that structural compliance and personal acceptance are different gates.

Write a three-part Structured Caption

MiniMax’s model card recommends three sections for precise control: Global Metadata, Vocal Details, and Arrangement. Global Metadata carries genre, tempo, key, emotional arc, scenario, and production profile. Vocal Details specifies the lead configuration; an instrumental bed states Instrumental — no vocals and names the lead instrument. Arrangement describes when instruments enter, change, intensify, and leave across named sections.

Local practice targets roughly 250–450 English words. Below about 120 words, the model tended to free-associate; past about 600, it tended to ignore the tail. These are local observations, while the model card supplies the official 5,000-token ceiling. The official MiniMax prompt guide recommends naming two or three instruments precisely and leaving the rest to the model. Local caption practice then gives every named instrument a lifecycle rather than leaving a flat shopping list.

Three stacked Structured Caption blocks showing what Global Metadata, Vocal Details, and Arrangement should carry and exclude
The useful constraint is not more adjectives; it is putting each instruction in the section that can carry it.

Separate soft targets from executable instructions

Requested BPM, key, instruments, and structure are soft targets. The model card says they provide generative control, not strict symbolic guarantees, so each needs a listening check. Section tags are the executable structural instructions in the local workflow. The supported tag set includes [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], and [Outro].

For an instrumental bed, the local skeleton is [Intro], [Instrumental], and [Outro]. Arrangement should name those sections so the caption and structural skeleton agree. A stray [Chorus] invited singing in local practice. That is an observed workflow warning, not a promise about every generated take.

Configure the ComfyUI node graph deliberately

The official ComfyUI tutorial describes max_duration as the target length and says the model supports about 300 seconds. Its shipped workflow template contains a discrepancy: the live widget is 60, while its own note says the default is 120. Neither value should be presented as one unqualified official default. Generation can also stop before the ceiling when the model emits an end-of-audio token.

A fixed seed reproduces the same song, while a changed seed produces another take, according to the tutorial. My local script used cfg_scale 1.7 and top_k 50, but the official tutorial and template do not explain those widgets’ semantics. I can record those local defaults without inventing official meanings. That makes one-variable tests more honest than simultaneous knob changes.

Tiled VAE decode reduces memory use, but the tutorial notes slower decoding and a small seam risk. It recommends disabling tiled decode on high-VRAM GPUs for best quality. This M1 Max workflow used tiles of 1536 with 64 overlap as its local memory-saving default. On Apple silicon, this setup used the fp16 DiT and pruned bf16 encoder; the local evidence warns that the CUDA-oriented int8_convrot variants should not be used on MPS.

Design narration beds and loops before delivery

A narration bed needs room for speech before mixing begins. Local guidance keeps 200 Hz–4 kHz sparse, removes the melodic hook, holds a flat emotional arc, and avoids risers, impacts, or gated-snare fills. The library rule is to audition each bed under actual narration rather than in isolation. How Instagram Reel scripts change across English, Portuguese, and Japanese provides the adjacent spoken-video context for that use.

A loop needs arrangement work before crossfading. The caption should end where the arrangement began. Then crossfade the tail into the head and listen across the seam twice. If the endpoints occupy different harmonic places, the local procedure treats the caption as the problem rather than extending the crossfade.

Check true peak before delivery

Fifteen of the 17 raw takes measured above 0 dBTP, with the loudest at +1.130 dBTP; yt-rock-opener measured exactly 0.000 dBTP, while ig-rock-launch measured −0.310 dBTP. Sixteen delivered files were attenuated to −1.200 dBTP; yt-boombap-city-night was attenuated to −1.400 dBTP. Those measurements describe only these files, not a general property of MiniMax Music 3.

The local release-master guidance targets about −14 LUFS integrated and a true peak at or below −1 dBTP. That is a workstation practice, not an official MiniMax or platform requirement. The important sequence is to measure, apply the intended level control, and verify the result before delivery.

What I think went wrong—and what I will change next

My interpretation is narrow: the process verified files and preserved inputs, but it did not place my taste decision early enough. Every caption had the required three sections, an instrumental Vocal Details block, and a timed Arrangement. Yet my own judgment is that the resulting tracks are not good enough to keep. A completed file and a personally accepted bed are different outcomes.

My interpretation of the long, prescriptive captions

The captions were also long and highly prescriptive. For example, ig-rock-launch dictated six timed sections and explicitly prohibited risers, impacts, gated-snare fills, crash-led transitions, and tom flourishes. I interpret that density as one place to test, not as a demonstrated cause of my verdict. This batch does not prove that detailed captions generally produce music someone will reject.

The next useful distinction is caption quality versus take quality. A fixed seed lets me compare one caption change against the same underlying take conditions. Once the caption survives that test, rolling seeds compares performances without rewriting the brief. Listening supplies the taste gate that structure and reproducibility cannot supply.

Second attempt: test one variable before a long render

I am going to make a second attempt soon. I will begin with one 30-second draft, hold the seed fixed, and change one caption axis. After freezing that caption, I will roll two or three seeds and choose by listening. Only then will I commit one seed to a long render with the intended duration plus about ten seconds.

The second attempt is a stated intention, not a promised result or schedule. Its purpose is to make each comparison interpretable. It does not assume that the next 17 tracks will pass my taste test.

Sources

Apply the five-line caption check to one 30-second draft

  1. Write Global Metadata with genre, tempo, key, emotional arc, scenario, and production profile.
  2. Write Vocal Details, or state Instrumental — no vocals and name the lead instrument.
  3. Bind Arrangement changes to the section tags used by the generation input.
  4. Name two or three instruments precisely, and give each one a lifecycle.
  5. Listen for the requested soft targets before committing the selected seed to a long render.

Keep reading