Scripts to experiment with six open-weights text-to-audio models:
Each model needs to be installed and run a bit differently, and in separate
Python venvs due to conflicting dependencies. See
linux_install.md or
macos_install.md. Running on MacOS in particular needs a
few patches and workarounds.
Once installed, the scripts provide:
*_single.py: generates N samples for a single prompt*_esc50.py: Re-runs the ESC-50 experiments from our paper, i.e. 100 samples each for the prompt "Sound of [label]" for each label in ESC-50. Note: generates 5000 audio files and will take a while! TODO: Not yet ported to all 6 models.*_param_count.py: Counts the number of parameters in each model. Added because I got frustrated trying to figure out how big the models were from their papers/documentation. TODO: Not yet ported to all 6 models.
The _single.py versions take a number of command line parameters. A few
examples:
# In venv_stableaudio:
python stableaudio_single.py --prompt "sound of a cat meowing" \
--samples 100 --duration 5
# In venv_ezaudio:
python ezaudio_single.py --ezaudio-repo /path/to/EzAudio \
--prompt "sound of a cat meowing" --samples 100 --duration 5
# In venv_tangoflux:
python tangoflux_single.py --prompt "sound of a cat meowing" \
--samples 100 --duration 5Our initial paper:
- Jonathan Morse, Azadeh Naderi, Swen Gaudl, Mark Cartwright, Amy K. Hoover, Mark J. Nelson (2025). Expressive range characterization of open text-to-audio models. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, pp. 91-98.