Audiyo's 1.49 GB Music3 download contains the compressed diffusion transformer, one part of the system that turns lyrics into songs. Music3 also needs a language model and an audio decoder, neither of which is included in that file.

The Python toolkit added Music3 support through several releases on September 11. Version 0.1.1 introduced the loading path, 0.2.0 added the compressed component, and 0.2.1 updated loading and required Diffusers 0.40 or newer. The documented workflow takes lyrics and a music description, then saves a WAV file. It requires installation, model downloads and a CUDA-capable GPU.

The upstream Diffusers guide puts its full pipeline at about 23 GB of GPU memory using bfloat16, a 16-bit format. Its slower 8 GB example combines automatic CPU offloading with group offloading, moving the language model's layers between system memory and the GPU. Those figures describe the upstream configuration; Audiyo's compressed setup needs its own measurements.

According to Audiyo, a 6 GB card can run its default compressed version. Its backend notes still call for a GPU check and describe some features as plans.

Audiyo's code is Apache-2.0. Its README incorrectly describes the model weights under a blanket reference to Stability AI's community license. The Music3 backend notes correctly identify MiniMax's separate terms. Commercial product interfaces must prominently display MiniMax-Music3. Prior written authorization is required when aggregate annual revenue from covered products and services, across the operator and its affiliates, exceeds $20 million. Its acceptable-use policy prohibits publicly sharing machine-generated information without clear and prominent disclosure.