From still to motion: the image-to-video workflow
Text-to-video gambles on composition and motion at once. Locking the frame first is cheaper, faster and far more controllable.
By Ideovex Team

A video render costs five times an image. Text-to-video asks a model to invent the composition and the motion in one shot, which means when you reject the result you are throwing away twenty-five credits. Splitting the problem is almost always the better trade.
The two-stage method
- Lock the frame. Generate stills until you have a first frame you would be happy to publish as a photograph. Iterate here — each attempt is five credits.
- Animate it. Send that still to an image-to-video model and describe only the movement. One video render, one result you already know the look of.
In the studio, the Animate action on any finished image carries the file straight into the video studio with the reference already attached, so you never download and re-upload.
Describe motion, not the picture
This is the part people get wrong. The model can already see your frame — repeating what is in it wastes the prompt. Write the camera move and the subject action:
- Camera —
slow push in,handheld drift right,locked off, no camera movement,crane up to reveal - Subject —
she turns her head toward the window,steam rises and curls,the flag snaps in the wind - Pace —
slow, deliberateorbrisk. Models over-animate by default; asking for restraint reads as more expensive.
One camera instruction plus one subject instruction is the sweet spot. Three competing motions in a five-second clip produce mush.
Pick the duration before you pick the model
Duration options follow the model — some offer four, six and eight seconds, others five and ten. Decide the cut length you need first, then filter the model gallery to models that can hit it. Generating ten seconds and trimming to four is a waste of credits and the extra seconds usually contain the drift and warping you did not want.
Where clips fall apart, and what to do
- Faces melt in the last second — shorten the clip, or push the subject further from camera so fewer pixels carry the identity.
- Hands and text warp — keep them out of the first frame, or crop past them. Motion models inherit whatever ambiguity the still already had.
- Nothing moves — your motion prompt was descriptive rather than instructional. Cut every adjective and leave the verbs.
- Everything moves — add
locked off cameraand give the subject exactly one action.
Working in batches
Once a look is dialled in, generate the whole sequence of first frames, then animate them in one sitting. Concurrency lets several renders run at once, and the queue panel shows position for anything waiting. Save the working recipe as a Studio Preset and the next episode is a five-minute job.


