Skip to content
Get started
GENERATION API CALLS
Audio Generation

Academia

This page is auto-generated from model configurations. Last updated: 2026-07-29.

This reference lists all available Academia audio generation models and their parameters. Use these parameter names when calling the Generation API.


Add one instrument stem on top of existing audio, matched to key, tempo, and groove. Build arrangements layer by layer.

Model ID: model_ace-step-1-5-edit-add-layer

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-edit-add-layer/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The existing track to layer onto.
trackNamestringYes---vocals, backing_vocals, drums, bass, guitar, keyboard, percussion, strings, synth, fx, brass, woodwindsThe stem to add on top of the source audio.
promptstringYes----Describe the new stem, e.g. “acoustic drum kit groove matching a fingerpicked folk guitar”.
repaintingStartnumberNo00--Where the new layer begins, in seconds.
repaintingEndnumberNo-1-1--Where the new layer ends, in seconds. Use -1 through to the end of the track.
thinkingbooleanNotrue---Lets the model plan the layer first for tighter musical coherence. On by default.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.
guidanceScalenumberNo7115-How strictly the result follows the prompt. Applies on the XL base edit tier.

Build full accompaniment around a partial track. Turn a bare vocal or solo instrument into a full arrangement (Vocal2BGM).

Model ID: model_ace-step-1-5-edit-complete-track

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-edit-complete-track/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The partial track to complete, such as an a cappella vocal or solo guitar.
completeTrackClassesstring_arrayYesdrums,bass,guitar--vocals, backing_vocals, drums, bass, guitar, keyboard, percussion, strings, synth, fx, brass, woodwindsWhich stems to generate around the source, e.g. drums, bass, and guitar.
promptstringYes----Overall style for the accompaniment, e.g. “rock style completion” or “warm lo-fi backing”.
thinkingbooleanNotrue---Lets the model plan the arrangement first for tighter musical coherence. On by default.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.
guidanceScalenumberNo7115-How strictly the result follows the prompt. Applies on the XL base edit tier.

Isolate a stem from a mixed track: vocals, drums, bass, guitar, and more. Uses the full-quality edit model with 50 diffusion steps.

Model ID: model_ace-step-1-5-edit-stem-extract

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-edit-stem-extract/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The mixed track to separate.
trackNamestringYes---vocals, backing_vocals, drums, bass, guitar, keyboard, percussion, strings, synth, fx, brass, woodwindsThe stem to isolate from the mix.
promptstringNo----Optional hint about the stem character, e.g. “female pop vocals”.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.
guidanceScalenumberNo7115-How strictly the result follows the prompt. Applies on the XL base edit tier.

Higher-fidelity cover and restyle. Preserve melody while changing genre, vocals, and arrangement, or borrow style from a reference audio.

Model ID: model_ace-step-1-5-quality-cover

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-quality-cover/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The song you want to cover or restyle.
referenceAudiofileNo----An optional track whose style you want to borrow: its genre, mood, and feel guide the result.
promptstringNo----Describe the style you want: genre, mood, instruments, vocal character, and production. For example, “acoustic folk, warm and mellow, female vocals.”
lyricsstringNo----Optional new or changed lyrics. Use section tags like [Verse] and [Chorus] to mark structure.
audioCoverStrengthnumberNo101-How closely the result follows the original song. Higher values stay faithful to the source; use low values (around 0.2) when you mainly want to borrow a style.
instrumentalbooleanNofalse---Creates a version without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
thinkingbooleanNotrue---Lets the model plan the track first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

Higher-fidelity region regeneration. Replace a section of an existing track guided by prompt and lyrics.

Model ID: model_ace-step-1-5-quality-repaint

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-quality-repaint/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The track that contains the section you want to regenerate.
promptstringNo----Describe the music for the new section: genre, mood, instruments, and tempo.
lyricsstringNo----Optional lyrics for the new section. Use section tags like [Verse] and [Chorus] to mark structure.
repaintingStartnumberNo00--Where the section to regenerate begins, in seconds.
repaintingEndnumberNo-1-1--Where the section to regenerate ends, in seconds. Use -1 to regenerate through to the end of the track.
instrumentalbooleanNofalse---Regenerates the section without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
thinkingbooleanNotrue---Lets the model plan the section first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

Higher-fidelity music generation with a 4B model and stronger planner. Lyrics, vocals, and structure tags; full tracks up to 10 minutes.

Model ID: model_ace-step-1-5-quality-text-to-music

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-quality-text-to-music/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringNo----Describe the music you want: genre, mood, instruments, and tempo. Either a prompt or lyrics is required.
lyricsstringNo----The lyrics to sing, with optional section tags like [Verse] and [Chorus]. Leave empty and turn on Instrumental for music without vocals. Either lyrics or a prompt is required.
instrumentalbooleanNofalse---Creates music without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
durationnumberNo-10600-How long the track should be, in seconds (10-600). Leave empty to let the model decide. Longer tracks cost more.
bpmnumberNo-30300-The tempo, in beats per minute. Leave empty to let the model choose.
keyscalestringNo----The musical key, such as “C Major” or “Am.” Leave empty to let the model choose.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
thinkingbooleanNotrue---Lets the model plan the track first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

Re-render or restyle an existing track with ACE-Step 1.5 turbo. Preserve melody while changing genre, vocals, and arrangement, or use a reference audio for style transfer.

Model ID: model_ace-step-1-5-turbo-cover

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-turbo-cover/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The song you want to cover or restyle.
referenceAudiofileNo----An optional track whose style you want to borrow: its genre, mood, and feel guide the result.
promptstringNo----Describe the style you want: genre, mood, instruments, vocal character, and production. For example, “acoustic folk, warm and mellow, female vocals.”
lyricsstringNo----Optional new or changed lyrics. Use section tags like [Verse] and [Chorus] to mark structure.
audioCoverStrengthnumberNo101-How closely the result follows the original song. Higher values stay faithful to the source; use low values (around 0.2) when you mainly want to borrow a style.
instrumentalbooleanNofalse---Creates a version without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
thinkingbooleanNotrue---Lets the model plan the track first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

Regenerate a time region in an existing track with ACE-Step 1.5 turbo. Replace a verse, chorus, or section with new music guided by prompt and lyrics.

Model ID: model_ace-step-1-5-turbo-repaint

Capabilities: audio2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-turbo-repaint/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
srcAudiofileYes----The track that contains the section you want to regenerate.
promptstringNo----Describe the music for the new section: genre, mood, instruments, and tempo.
lyricsstringNo----Optional lyrics for the new section. Use section tags like [Verse] and [Chorus] to mark structure.
repaintingStartnumberNo00--Where the section to regenerate begins, in seconds.
repaintingEndnumberNo-1-1--Where the section to regenerate ends, in seconds. Use -1 to regenerate through to the end of the track.
instrumentalbooleanNofalse---Regenerates the section without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
audioFormatstringNomp3--mp3, wav, flacThe audio file type you get back.
thinkingbooleanNotrue---Lets the model plan the section first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

Open-source music generation with lyrics, vocals, and structure tags like [Verse] and [Chorus]. Full tracks up to 10 minutes with ACE-Step v1.5 turbo.

Model ID: model_ace-step-1-5-turbo-text-to-music

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_ace-step-1-5-turbo-text-to-music/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringNo----Describe the music you want: genre, mood, instruments, and tempo. Either a prompt or lyrics is required.
lyricsstringNo----The lyrics to sing, with optional section tags like [Verse] and [Chorus]. Leave empty and turn on Instrumental for music without vocals. Either lyrics or a prompt is required.
instrumentalbooleanNofalse---Creates music without vocals.
vocalLanguagestringNounknown--ar, az, bg, bn, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, ht, hu, id, is, it, ja, ko, la, lt, ms, ne, nl, no, pa, pl, pt, ro, ru, sa, sk, sr, sv, sw, ta, te, th, tl, tr, uk, ur, vi, yue, zh, unknownThe language for the vocals. Leave as Auto to detect it automatically.
durationnumberNo-10600-How long the track should be, in seconds (10-600). Leave empty to let the model decide. Longer tracks cost more.
bpmnumberNo-30300-The tempo, in beats per minute. Leave empty to let the model choose.
keyscalestringNo----The musical key, such as “C Major” or “Am.” Leave empty to let the model choose.
numOutputsnumberNo114-How many audio variations to generate (1-4). More outputs cost more.
thinkingbooleanNotrue---Lets the model plan the track first, which improves how closely it follows your prompt. On by default.
seednumberNo-0--A number that makes the first result repeatable; any extra outputs use random seeds. Leave empty for a fresh result each time.

High-quality voice cloning TTS at 48kHz from text and a reference audio clip.

Model ID: model_lux-tts

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_lux-tts/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Text to convert to speech.
audiofileYes----Reference audio for voice cloning.
guidanceScalenumberNo3010-Higher values increase adherence to the reference voice.
numInferenceStepsnumberNo4116-Number of flow-matching inference steps.
maxRefLengthnumberNo5115-Maximum reference audio duration used for voice encoding (seconds).
seednumberNo-02147483647-Seed for reproducible outputs.

MMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.

Model ID: model_mm-audio-2-t2a

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_mm-audio-2-t2a/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
promptstringYes----Text prompt for generated audio
negativePromptstringNo----Negative prompt to avoid certain sounds
durationnumberNo8130-Output duration in seconds.
numStepsnumberNo25450-The number of steps to generate the audio for
cfgStrengthnumberNo4.5120-Higher values will keep output closer to the prompt
maskAwayClipbooleanNofalse---Mask away certain sounds in the audio
seednumberNo-065535-Random seed for reproducible generation

Lighter Tada voice cloning text-to-speech variant with multilingual support.

Model ID: model_tada-1b-text-to-speech

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_tada-1b-text-to-speech/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
audiofileYes----Reference audio for voice cloning.
promptstringYes----Text to synthesize with the reference voice.
transcriptstringNo----Transcript of the reference audio. Required for non-English references.
languagestringNoen--en, ar, ch, de, es, fr, it, ja, pl, ptLanguage used for text alignment.
numInferenceStepsnumberNo20150-Number of ODE solver steps for acoustic generation.
speedUpFactornumberNo10.52-Values > 1 speed up and values < 1 slow down speech.
temperaturenumberNo0.602-Sampling temperature for text token generation.
topPnumberNo0.901-Top-p nucleus sampling value.
repetitionPenaltynumberNo1.112-Penalty applied to repeated tokens.
acousticCfgScalenumberNo1.6010-Classifier-free guidance scale for acoustic generation.
noiseTemperaturenumberNo0.902-Temperature for diffusion noise during flow matching.
numExtraStepsnumberNo0050-Additional autoregressive steps for continuation.

Voice cloning text-to-speech with multilingual alignment and expressive controls.

Model ID: model_tada-3b-text-to-speech

Capabilities: txt2audio

LLM Markdown: https://app.scenario.com/api/models/model_tada-3b-text-to-speech/markdown

ParameterTypeRequiredDefaultMinMaxAllowed ValuesDescription
audiofileYes----Reference audio for voice cloning.
promptstringYes----Text to synthesize with the reference voice.
transcriptstringNo----Transcript of the reference audio. Required for non-English references.
languagestringNoen--en, ar, ch, de, es, fr, it, ja, pl, ptLanguage used for text alignment.
numInferenceStepsnumberNo20150-Number of ODE solver steps for acoustic generation.
speedUpFactornumberNo10.52-Values > 1 speed up and values < 1 slow down speech.
temperaturenumberNo0.602-Sampling temperature for text token generation.
topPnumberNo0.901-Top-p nucleus sampling value.
repetitionPenaltynumberNo1.112-Penalty applied to repeated tokens.
acousticCfgScalenumberNo1.6010-Classifier-free guidance scale for acoustic generation.
noiseTemperaturenumberNo0.902-Temperature for diffusion noise during flow matching.
numExtraStepsnumberNo0050-Additional autoregressive steps for continuation.