Prompt Lab: testing prompt changes on purpose
Structured experiments that isolate one prompt variable at a time, so you learn what a phrase actually does instead of collecting anecdotes.
Why a lab at all
Generative models are non-deterministic. Run the same prompt twice and you get two different songs, which means a single generation can never tell you whether a change helped. Most prompt advice on the internet is exactly that: one good result, remembered as a rule.
The Lab is our answer. Each study fixes everything except one phrase, runs a set number of generations on each side, scores them against a rubric written before listening, and records a verdict — including "no detectable difference", which is a perfectly good result and the one people forget to publish.
The eight-generation test
The house method. Eight generations total: four with the control prompt, four with the variant. Eight is a pragmatic number — enough that a single lucky take cannot carry the result, small enough that you will actually finish the test in one sitting.
- Write the hypothesis first. One sentence, falsifiable. "Naming a reverb character will make the vocal sit further back in the mix."
- Build the prompt pair. Identical strings except for the phrase under test. Use Prompt Diff to confirm nothing else drifted.
- Fix the lyrics. Same lyric sheet on both sides, or instrumental on both sides. Lyrics change arrangement more than people expect.
- Generate four per side, alternating A, B, A, B so that any drift over the session hits both arms equally.
- Write the listening rubric before you listen. Two or three criteria, each scored 0–2. Deciding what counts after hearing the results is how you fool yourself.
- Score blind if you can. Rename files to numbers and shuffle them.
- Record the verdict in one of four buckets: clear effect, weak effect, no effect, or unstable (the variant works sometimes and backfires other times).
- Keep the whole set. Losing takes are the evidence.
Score sheet template
Copy this table into your notes at the start of a study and fill it in as you go. The values below are placeholders showing the shape of a completed sheet — they are not results from any run.
When you read your filled-in sheet: a difference worth acting on usually shows up as a gap of at least one full point in arm means, plus a consistent direction across individual takes. Anything smaller is noise at this sample size.
| Take | Arm | Criterion 1 (0–2) | Criterion 2 (0–2) | Criterion 3 (0–2) | Total | Note |
|---|---|---|---|---|---|---|
| 01 | A — control | — | — | — | — | template row |
| 02 | B — variant | — | — | — | — | template row |
| 03 | A — control | — | — | — | — | template row |
| 04 | B — variant | — | — | — | — | template row |
| 05–08 | alternating | — | — | — | — | template rows |
| Arm mean | A vs B | — | — | — | — | fill after scoring |
Published studies
- Vocal space: does naming reverb character move the vocal back? — a production-language study.
- Chorus lift: does an explicit lift instruction raise chorus energy? — an arrangement study.
Both studies publish the full design — hypothesis, prompt pair, rubric and verdict framework — so you can rerun them on your own account and compare notes.
Run your own
Good candidate variables share three properties: you can express the change in a single phrase, you can hear the effect without reference gear, and you actually care about the answer. Bad candidates are vague adjectives, whole-prompt rewrites, and anything you cannot describe to another person in one sentence.
Start with a variable you expect to matter. Confirming something obvious teaches you the method cheaply, and gives you a baseline feel for how much variation a single arm produces.
- Tempo language: "120 BPM" versus "uptempo".
- Instrument naming: one hero instrument versus a list of five.
- Vocal descriptors: type only versus type plus delivery.
- Negative phrasing: does excluding something work better than not mentioning it?
- Section tags: bare lyrics versus tagged sections.