Models
Every model on Dolly, included on every plan. Each one shows its exact credit cost on the button before you run it.
Video (15)
Seedance 2.5Seedance 2.5 is ByteDance's newest video flagship, built for long single shots with reference control and an optional generated soundtrack.
Seedance 2.0Seedance 2.0 is ByteDance's established video flagship and the only model here that renders at 4k.
Seedance 2.0 MiniSeedance 2.0 Mini is the cheap, quick end of the Seedance family, with the same frame and reference controls as its bigger siblings.
Seedance 2.0 FastSeedance 2.0 Fast is the speed tier of the Seedance line, trading top resolution for turnaround while keeping frame control.
Seedance 1.5 ProSeedance 1.5 Pro is the previous generation Seedance Pro, still one of the cheapest ways to get a controlled clip with optional audio.
Kling 3.0Kling 3.0 is Kuaishou's latest video model and the strongest here for physical motion, with 4k output and first and last frame control.
Kling 3.0 TurboKling 3.0 Turbo is the faster, cheaper Kling tier, keeping the family's motion quality with first frame control only.
Veo 3.1Veo 3.1 is Google's quality tier for video and the model here with the best native audio, generating speech and sound with the picture.
Veo 3.1 FastVeo 3.1 Fast is the Veo speed tier, keeping native audio and adding reference image support that the quality tier does not have.
Veo 3.1 LiteVeo 3.1 Lite is the cheapest way to get Veo's native audio, with reference support and the full resolution ladder.
Wan 2.7Wan 2.7 is Alibaba's video model, notable for taking a reference image together with a first frame and for exposing a seed and a negative prompt.
Hailuo 2.3Hailuo 2.3 is MiniMax's image to video model and it requires a starting image, which is exactly what makes it good at animating a still you already have.
Grok Imagine 1.5Grok Imagine 1.5 is xAI's video model and the cheapest clip on Dolly, taking text or a single starting image.
HappyHorse 1.1HappyHorse 1.1 is Alibaba's video model with the widest aspect ratio menu here and a seed you can pin.
Gemini OmniGemini Omni is Google's reference driven video model, taking up to 4 images in and generating picture with native audio.Image (7)
Nano Banana ProNano Banana Pro is Google's top image model, with reference support, the widest aspect ratio menu here and output up to 4K.
GPT Image 2GPT Image 2 is OpenAI's image model and the best here at following a complicated written instruction exactly.
Nano Banana 2Nano Banana 2 is Google's fast image model, with the same wide aspect menu and 4K ceiling as the Pro tier at a lower price.
Seedream 5 ProSeedream 5 Pro is ByteDance's image flagship, strongest here on fine detail and material texture.
Seedream 4.5Seedream 4.5 is the value tier of the Seedream line and the only image model here that starts at 2K, with 4K at the same price.
Flux 2 ProFlux 2 Pro is Black Forest Labs' precise, editorial image model and the cheapest still on Dolly.
Flux 2 FlexFlux 2 Flex is the flexible Black Forest Labs tier, trading some of Pro's precision for a wider range of looks.Voice (6)
ElevenLabs Multilingual V2ElevenLabs Multilingual V2 is the classic ElevenLabs voice model, with the largest voice library on Dolly and consistently lifelike delivery.ElevenLabs V3ElevenLabs V3 is the expressive ElevenLabs model, capable of multi voice dialogue in a single generation.Gemini 2.5 Pro TTSGemini 2.5 Pro TTS is Google's studio quality speech model, with a curated voice set and multi speaker scene support.KokoroKokoro is a fast, open source speech model with twenty American voices and the lowest price per word on Dolly.MiniMax Speech-02 HDMiniMax Speech-02 HD is MiniMax's studio voice model, with a compact character voice set and broad language coverage.DiaDia is a dialogue first speech model: you write a script with speaker tags and it renders the whole exchange in one take.
Music (4)
Suno V5.5Suno V5.5 is Suno's flagship music model, returning two takes per generation and letting you set the track length directly.Suno V5Suno V5 is the fast, expressive Suno tier, returning two takes per generation for the same flat price.Suno V4.5+Suno V4.5+ is the long form Suno tier, built for full length songs rather than short cues, with two takes per generation.Eleven MusicEleven Music is ElevenLabs' music model, returning one polished, mix ready take with a length you set.