Write the operation, not the picture. A prompt that names a specific action and how much of it, like "rotate the camera 180 degrees", gets obeyed. A prompt that describes the result you are imagining, like "a low angle looking up at him", mostly does not. We tested this across 24 generations where the only thing that changed was the wording, and the gap was not subtle: 4 out of 4 against 0 out of 4.
We had a real problem to solve. Asking for a new view of an existing location kept returning the same camera position, and we could not tell whether the model was ignoring the text or the reference image was overpowering it. So we ran the same request 24 times across two locations, changing only the phrasing.
Every cell is four generations. Obeyed means the camera actually moved where the words asked it to:
| What you type | Camera obeyed |
|---|---|
| "rotate the camera 180 degrees" | 4 of 4 |
| "an overhead drone shot" | 3 of 4 |
| "from behind the tower" | 3 of 4 |
| "a steep high angle, looking down" | 2 of 4 |
| "orbit about 30 degrees to the right" | 1 of 4 |
| "a low angle looking up" | 0 of 4 |
Read the two ends of that table. The winners name a camera operation and give it an unambiguous size. The losers describe the photograph you are hoping to receive. Same model, same reference image, same everything else. Only the words changed.
A phrase like "a steep high angle" is a description of an outcome. It assumes the model will work backwards from the picture in your head to the camera move that produces it. Sometimes it does. Often it gives you the subject you named and leaves the camera where it was, because nothing in the sentence was an instruction to move anything.
"Rotate the camera 180 degrees" cannot be satisfied without moving the camera. There is no version of that sentence the model can answer while doing nothing. That is the whole trick, and it generalises past cameras: prefer the sentence that is false unless the change happens.
One honest caveat on our own numbers, because it matters if you are about to trust them. The 30 degree orbit scored 1 of 4, but 30 degrees is a small move and our judgement of whether a subtle rotation happened is softer than our judgement of a 180. A 180 succeeding does not prove a 30 degree turn was refused. The high and low angle rows are the strong evidence; the orbit row is the weak one.
The rule is portable. Take the thing you want changed, and write the change as an operation with a magnitude rather than as a description of the finished frame.
| Instead of | Write |
|---|---|
| "a wider shot of the garden" | "zoom the camera out until the whole pond is in frame" |
| "make her look worried" | "change her expression to worried, eyebrows raised" |
| "a more dramatic angle" | "rotate the camera 90 degrees to the right" |
| "the same scene but at night" | "change the time of day to night, keep everything else" |
Notice the right column is longer and duller. That is the point. Prompts are not writing, and the adjectives that make a sentence read well are exactly the words that give a model room to do nothing.
Most prompting advice is about the first generation. Most of your actual time goes on the second, third and fourth, when something is nearly right and one element is wrong. That is a different skill, and the common mistake is to rewrite the whole prompt.
Rewriting everything re-rolls everything. If you liked the composition and hated the pose, a fresh description of the whole frame gives the model licence to move the parts you were happy with. Name the one change and say to keep the rest. This is the same rule as before: an operation on what exists, not a description of a new picture.
The second habit worth building is to stop after two. If a specific, operation-shaped instruction has failed twice, a third attempt at the same sentence is not iteration, it is hoping. Change the approach, not the wording: fix it as a still, or split the moment into two keyframes so the thing you cannot describe becomes a thing you can place.
Prompting is how you control what a frame contains. It is a poor way to control what happens over time, because time is where a text description is vaguest and a model has the most room.
This is one shot from a three shot scene, ukiyo-e style. A dragon at a koi pond catches a fish and burps a fireball:
None of that motion was prompted. Across the whole scene, nine still frames were authored, two cuts were placed between them, and the model rendered the 474 frames that came out. Every shot came back usable on the first render, with no retries, because there was nothing left for it to guess about where things ended up.
That is the real answer to prompts that will not behave. Past a certain point you stop describing motion and start placing its endpoints. We wrote up how that works in how to add keyframes and cuts to an AI video.
Nima is built for the operation-shaped way of working: you describe a change to something that already exists, and the thing you were happy with stays put.
The free plan includes finished films with their keyframes intact, so you can test operation-shaped prompts against a scene that already works. Open Nima free.
Usually because the prompt described a result rather than an action. A model can satisfy "a dramatic low angle" by giving you the subject and leaving the camera alone, since nothing in that phrase is an instruction to move it. Rewrite it as an operation with a size, like "rotate the camera 90 degrees to the right", and there is no way to answer it without doing the thing.
Name the single change and say to keep everything else. Rewriting the whole description re-rolls the whole image, including the parts that were already right. Work in tools that keep the previous version so a regeneration you dislike costs nothing.
Twice, on the same instruction. If a specific, operation-shaped prompt has failed twice, a third try at the same sentence is hoping rather than iterating. Change the method instead: edit it as a still, or split the moment into two keyframes.
Longer is not the variable. Specific is. In our test the winning phrasings were short and named an operation and a magnitude, while some of the losing ones were longer and more descriptive. Adding adjectives adds room for the model to do nothing.
No. The pattern is that instructions which cannot be satisfied without a change get obeyed, and descriptions of a hoped-for result often do not. That holds for expressions, time of day, position and framing. Cameras are just where we measured it.
Nima lets you fix a frame by naming one change, and keeps the version you already liked. Try it free.