Prompt Lab

Prompt Lab: testing prompt changes on purpose

Structured experiments that isolate one prompt variable at a time, so you learn what a phrase actually does instead of collecting anecdotes.

Why a lab at all

Generative models are non-deterministic. Run the same prompt twice and you get two different songs, which means a single generation can never tell you whether a change helped. Most prompt advice on the internet is exactly that: one good result, remembered as a rule.

The Lab is our answer. Each study fixes everything except one phrase, runs a set number of generations on each side, scores them against a rubric written before listening, and records a verdict — including "no detectable difference", which is a perfectly good result and the one people forget to publish.

The eight-generation test

The house method. Eight generations total: four with the control prompt, four with the variant. Eight is a pragmatic number — enough that a single lucky take cannot carry the result, small enough that you will actually finish the test in one sitting.

Eight generations is a craft heuristic, not a statistical test. It is enough to catch obvious effects and to stop you generalising from one happy accident. It cannot detect small differences, and it will not survive a statistician's scrutiny. Treat every verdict as a working note.
  • Write the hypothesis first. One sentence, falsifiable. "Naming a reverb character will make the vocal sit further back in the mix."
  • Build the prompt pair. Identical strings except for the phrase under test. Use Prompt Diff to confirm nothing else drifted.
  • Fix the lyrics. Same lyric sheet on both sides, or instrumental on both sides. Lyrics change arrangement more than people expect.
  • Generate four per side, alternating A, B, A, B so that any drift over the session hits both arms equally.
  • Write the listening rubric before you listen. Two or three criteria, each scored 0–2. Deciding what counts after hearing the results is how you fool yourself.
  • Score blind if you can. Rename files to numbers and shuffle them.
  • Record the verdict in one of four buckets: clear effect, weak effect, no effect, or unstable (the variant works sometimes and backfires other times).
  • Keep the whole set. Losing takes are the evidence.

Score sheet template

Copy this table into your notes at the start of a study and fill it in as you go. The values below are placeholders showing the shape of a completed sheet — they are not results from any run.

When you read your filled-in sheet: a difference worth acting on usually shows up as a gap of at least one full point in arm means, plus a consistent direction across individual takes. Anything smaller is noise at this sample size.

TakeArmCriterion 1 (0–2)Criterion 2 (0–2)Criterion 3 (0–2)TotalNote
01A — controltemplate row
02B — varianttemplate row
03A — controltemplate row
04B — varianttemplate row
05–08alternatingtemplate rows
Arm meanA vs Bfill after scoring

Published studies

Both studies publish the full design — hypothesis, prompt pair, rubric and verdict framework — so you can rerun them on your own account and compare notes.

Run your own

Good candidate variables share three properties: you can express the change in a single phrase, you can hear the effect without reference gear, and you actually care about the answer. Bad candidates are vague adjectives, whole-prompt rewrites, and anything you cannot describe to another person in one sentence.

Start with a variable you expect to matter. Confirming something obvious teaches you the method cheaply, and gives you a baseline feel for how much variation a single arm produces.

  • Tempo language: "120 BPM" versus "uptempo".
  • Instrument naming: one hero instrument versus a list of five.
  • Vocal descriptors: type only versus type plus delivery.
  • Negative phrasing: does excluding something work better than not mentioning it?
  • Section tags: bare lyrics versus tagged sections.
Building the prompt pair is faster in the Prompt Builder — construct the control, copy it, change one pillar, and diff the two strings before you generate.

FAQ

Why eight and not thirty?
Thirty would be better and almost nobody completes it. Eight fits one focused session, survives a single outlier, and is honest about being a craft heuristic rather than a statistical test.
Do you publish raw audio from the studies?
No. We publish the design — hypothesis, prompt pair, rubric, verdict framework — so you can run it yourself. Your account, your model version and your lyrics all affect the outcome, so your own run is worth more than our files.
What if the result is 'no effect'?
That is a useful finding and we publish it. Knowing a phrase does nothing saves you characters in every future prompt.
Can I submit a study?
Yes — send the design and your score sheet through the contact page. We are interested in the method being followed, not in whether the result flatters anyone.