Playcut AI models: specs, limits and credit pricing
Playcut gives you one API across models from Google, xAI, ByteDance and Alibaba — one API key, one credit balance, one async task shape. Switching from Veo 3.1 to Seedance 2.0 is a model field change, not an integration.
This page is the index: three image models, seven video models, two music models and three voice providers, with the numbers you actually need to pick one. Every model has its own page with per-mode detail and known failure modes.
Image models
Section titled “Image models”Three models, one order of magnitude apart in price. All three do text-to-image and reference-to-image (editing); only Nano Banana Pro does multi-region edit.
| Model | Best for | Modes | Max resolution | Credits | Page |
|---|---|---|---|---|---|
Nano Banana Progemini-3-pro-image-preview | Highest fidelity, text rendering, multi-region editing | text-to-image, reference-to-image, multi-region edit | 4K | 67 @ 1K/2K · 84 @ 4K · 25 flat per multi-region edit | Nano Banana Pro |
Nano Banana 2gemini-3.1-flash-image-preview | The workhorse — near-Pro quality at about half the price | text-to-image, reference-to-image | 4K | 34 @ 1K/2K · 43 @ 4K | Nano Banana 2 |
Grok Imagine (Image)grok-imagine-image | Cheap bulk generation, thumbnails, style exploration | text-to-image, reference-to-image | 2K (no 4K) | 10 text-to-image · 20 reference-to-image | Grok Imagine Image |
Both Nano Banana models share the same ten aspect ratios (1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9). Grok Imagine Image is the only image model with an auto ratio and the only one offering 2:1, 1:2, 19.5:9, 9:19.5, 20:9 and 9:20 — but it stops at 2K.
Video models
Section titled “Video models”Seven models. They differ far more than the image models do — in modes, in maximum duration, and by 24× in price per second (10 credits/s on Seedance 2.0 at 480p up to 240 credits/s on Veo 3.1 at 4K).
Capabilities
Section titled “Capabilities”| Model | Best for | Modes | Aspect ratios | Max resolution | Duration |
|---|---|---|---|---|---|
Veo 3.1veo-3.1-generate-preview | Highest quality, native audio, 4K, the only model with all five modes | text-to-video, image-to-video, reference-to-video, interpolation, extension | 16:9, 9:16 only | 4K | 4, 6 or 8 s (enum). Must be 8 at 1080p, 4K, with references, or for extension. Extension adds 7 s per call, up to 20 times |
Grok Imagine (classic)grok-imagine-video | Cheapest video per second; widest ratio list | text-to-video, image-to-video, reference-to-video, extension | 7 ratios (16:9, 4:3, 1:1, 9:16, 3:4, 3:2, 2:3) | 720p | 1–15 s text-to-video · 1–10 s for other modes · default 6 s |
Grok Imagine 1.5grok-imagine-video-1.5-preview | Better motion than classic, still cheap | image-to-video (text-to-video shipped by Playcut — see caveat) | inherited from classic, unverified | 720p | 1–15 s text-to-video · 1–10 s image-to-video |
Seedance 2.0seedance-2.0 | Best capability-per-credit; the only non-Veo 4K | text-to-video, image-to-video, reference-to-video (+ audio ref), transition | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 4K | 4–15 s (integer), default 5 s |
Gemini Omni Flashgemini-omni-flash-preview | Fixed-length 10 s animate clips from a still | image-to-video only on Playcut | output follows the source image | 720p | Fixed 10 s (model actually returns 3–10 s; no duration parameter exists) |
HappyHorse 1.1happyhorse-1.1 | Reference-heavy shots — up to 9 reference images; video edit with reference swaps | text-to-video, image-to-video, reference-to-video, video edit | all 9 the provider accepts | 1080P | 3–15 s (integer); edit follows the source (≤15 s) |
Wan 2.7wan-2.7 | Lip-sync from driving audio, and the only video-edit model | image-to-video (audio-driven), reference-to-video, video-edit | 16:9, 9:16, 1:1, 4:3, 3:4 — image-to-video takes no ratio at all | 720p | 4–15 s (video-edit: provider range 2–10 s) |
Pricing and reference limits
Section titled “Pricing and reference limits”| Model | Credits per second | Max reference images | Page |
|---|---|---|---|
| Veo 3.1 | 160 @ 720p/1080p · 240 @ 4K | 3 (reference) · 1 (animate) · 2 (interpolation) · 1 video (extension) | Veo 3.1 |
| Grok Imagine (classic) | 20 @ 480p/720p | 7 | Grok Imagine Video |
| Grok Imagine 1.5 | 32 @ 480p/720p | 0 — 1.5 rejects reference_images with a 400 | Grok Imagine Video |
| Seedance 2.0 | 10 @ 480p · 21 @ 720p · 60 @ 1080p · 108 @ 4K | 7 on Playcut (provider allows 9) | Seedance 2.0 |
| Gemini Omni Flash | 40 @ 720p (400 credits for the fixed 10 s) | 0 | Gemini Omni Flash |
| HappyHorse 1.1 | 56 @ 720P · 72 @ 1080P | 9 (reference) · exactly 1 first frame (animate) | HappyHorse 1.1 |
| Wan 2.7 | 40 @ 720p | 5 combined (reference) · 1 first frame (animate) · 1 source video (edit) | Wan 2.7 |
Music models
Section titled “Music models”Two Lyria 3 tiers from Google. Both are text-to-music and both accept up to 10 reference images as multimodal mood input — there is no audio-reference input.
| Model | Best for | Output | Duration | Max references | Credits | Page |
|---|---|---|---|---|---|---|
Lyria 3 Cliplyria-3-clip-preview | Loops, stings, background beds, fast iteration | 44.1 kHz stereo MP3 | Fixed 30 s — no duration parameter | 10 images | 20 per song | Lyria 3 |
Lyria 3 Prolyria-3-pro-preview | Full arranged tracks with verse/chorus structure | 44.1 kHz stereo, MP3 | ~2 minutes, influenced by the prompt only | 10 images | 40 per song | Lyria 3 |
Neither tier exposes a duration parameter. On Pro, arrangement is steered with [Verse] / [Chorus] / [Bridge] section markers in the prompt.
Voice models
Section titled “Voice models”Three providers behind one generate-tts tool. TTS is billed per 1,000 characters of input, rounded up, minimum one block.
| Model | Provider | Best for | Provider input cap | Credits | Page |
|---|---|---|---|---|---|
qwen3-tts-instruct-flash-2026-01-26 | Qwen (default) | Everyday synthesis, lowest cost | 600 characters | 15 / 1,000 chars | TTS Voices |
qwen3-tts-vd-2026-01-26 | Qwen | Designing a new voice from a text description | not documented | 20 flat per voice design | TTS Voices |
qwen3-tts-vc-2026-01-22 | Qwen | Cloning a voice from an audio sample | sample 5–60 s, ≤10 MB, ≥24 kHz, WAV/MP3/M4A | 5 per clone | TTS Voices |
gemini-3.1-flash-tts-preview | Gemini | Best stylistic control and language coverage | 32k-token session context (Playcut DTO caps at 5,000 chars) | 25 / 1,000 chars | TTS Voices |
How to choose
Section titled “How to choose”Cheapest per output
- Image — Grok Imagine Image, 10 credits.
- Video — Seedance 2.0 at 480p, 10 credits per second. At 720p, Grok Imagine classic (20) and Seedance (21) are effectively tied.
- Voice — Qwen, 15 credits per 1,000 characters.
- Music — Lyria 3 Clip, 20 credits per 30-second song.
Highest quality
- Image — Nano Banana Pro, the only image model with multi-region editing.
- Video — Veo 3.1, the only model with native synchronized audio and the only one reaching 4K.
- Voice — Gemini, the priciest TTS we sell at 25 credits per 1,000 characters.
Longest single clip
- Seedance 2.0, HappyHorse 1.1 and Wan 2.7 all reach 15 seconds in one generation. Both Grok Imagine models also reach 15 seconds, but only on text-to-video — their image- and reference-driven modes cap at 10 seconds.
- Veo 3.1 caps a single clip at 8 seconds but is the only model that can extend — 7 seconds per call, up to 20 times.
Needs lip-sync from an audio track
- Wan 2.7 only. It is the sole model that takes a driving
audioAssetIdon image-to-video.
Needs to edit an existing video
- Wan 2.7 only. Every other model 400-rejects video-edit on Playcut.
Needs many reference images
- Video — HappyHorse 1.1 (9) then Seedance 2.0 (7) and Grok Imagine classic (7).
- Image — a Nano Banana model. Avoid Grok Imagine Image for multi-reference work; only the first reference is sent.
Needs vertical or square video
- Anything except Veo 3.1, which is
16:9and9:16only. The widest video ratio list is Grok Imagine’s 7 (16:9,4:3,1:1,9:16,3:4,3:2,2:3); Seedance 2.0 and HappyHorse 1.1 offer 6 each and are the only video models with21:9.
Needs 4K
- Image — Nano Banana Pro (84 credits) or Nano Banana 2 (43 credits).
- Video — Veo 3.1 only, at 240 credits per second.
Cross-model references
Section titled “Cross-model references”Actor engine tiers
Section titled “Actor engine tiers”The actor-act pipeline exposes four engine tiers instead of raw model ids. Each maps onto a video model above, and each adds a flat 15 credit scene charge:
| Tier | Underlying model | Credits per second |
|---|---|---|
fastest (default) | Grok Imagine classic, reference-to-video | 7 |
fast | Grok Imagine 1.5, image-to-video | 17 |
plus | HappyHorse 1.1, reference-to-video | 30–38 |
pro | Gemini Omni Flash | 35 |
The tier label is the only engine identifier on the wire — the underlying provider is an implementation detail and can change.
Next steps
Section titled “Next steps”- Quickstart — first generation in under 5 minutes
- API Reference — every tool and its schema
- Rate Limits — concurrency and throughput per plan