All writing
Tutorial 2 min readAug 17, 2026

From still to motion: the image-to-video workflow

Text-to-video gambles on composition and motion at once. Locking the frame first is cheaper, faster and far more controllable.

By Ideovex Team

A video render costs five times an image. Text-to-video asks a model to invent the composition and the motion in one shot, which means when you reject the result you are throwing away twenty-five credits. Splitting the problem is almost always the better trade.

The two-stage method

  1. Lock the frame. Generate stills until you have a first frame you would be happy to publish as a photograph. Iterate here — each attempt is five credits.
  2. Animate it. Send that still to an image-to-video model and describe only the movement. One video render, one result you already know the look of.

In the studio, the Animate action on any finished image carries the file straight into the video studio with the reference already attached, so you never download and re-upload.

Describe motion, not the picture

This is the part people get wrong. The model can already see your frame — repeating what is in it wastes the prompt. Write the camera move and the subject action:

  • Camera — slow push in, handheld drift right, locked off, no camera movement, crane up to reveal
  • Subject — she turns her head toward the window, steam rises and curls, the flag snaps in the wind
  • Pace — slow, deliberate or brisk. Models over-animate by default; asking for restraint reads as more expensive.

One camera instruction plus one subject instruction is the sweet spot. Three competing motions in a five-second clip produce mush.

Pick the duration before you pick the model

Duration options follow the model — some offer four, six and eight seconds, others five and ten. Decide the cut length you need first, then filter the model gallery to models that can hit it. Generating ten seconds and trimming to four is a waste of credits and the extra seconds usually contain the drift and warping you did not want.

Where clips fall apart, and what to do

  • Faces melt in the last second — shorten the clip, or push the subject further from camera so fewer pixels carry the identity.
  • Hands and text warp — keep them out of the first frame, or crop past them. Motion models inherit whatever ambiguity the still already had.
  • Nothing moves — your motion prompt was descriptive rather than instructional. Cut every adjective and leave the verbs.
  • Everything moves — add locked off camera and give the subject exactly one action.

Working in batches

Once a look is dialled in, generate the whole sequence of first frames, then animate them in one sitting. Concurrency lets several renders run at once, and the queue panel shows position for anything waiting. Save the working recipe as a Studio Preset and the next episode is a five-minute job.

#video#workflow#image-to-video

Go try it.

Everything in this article works in the studio right now, with free credits on signup.

110+

AI models

Image, video & music — one prompt bar

22 / day

Free credits

Plus 0-credit models. No card to start.

Auto-refund

On failed renders

Credits reserved on start, refunded if it fails

One balance

For everything

Image, video and music on the same credits