AI image model benchmark
The same product brief through eleven image models, one attempt each, with the price, the time and the pixels you actually get back.
The run
- Models attempted
- 11
- Delivered a file
- 8
- Failed outright
- 3
- Whole run cost
- 92 cr ≈ 26¢
- Fastest
- Flux 2 Dev · 9.5s
- Run on
- 12 September 2026
A still is the cheapest thing to test properly, so there is no excuse for guessing. This is one product-photography brief — a hard one, with a reflective surface and directional light — sent once to eleven image engines spanning every price tier in the studio, from one credit to twenty.
What comes out is not a beauty contest. It is a price list with evidence attached: which engines are fast, which quietly return a bigger frame than you asked for, which cost twenty times more for a smaller file, and which do not work at all.
The brief, verbatim
A matte black insulated water bottle standing on wet slate, soft overcast light from the left, water droplets on the metal, faint reflection under the base, product photography, shallow depth of field.
Settings: 1024x1024, one image, no negative prompt. One attempt per model, run one after another, no re-rolls and no per-model prompt tuning.
The numbers
What each engine charged and how long it took.
Wall-clock from submit to a finished file on one account, single run. Credits are the studio's list price, which is what the account was charged.
| Model | Price | Time | Delivered | Result |
|---|---|---|---|---|
| Flux 2 Devwithdrawn | 1 cr | 9.5s | 1024×1024 · 0.04 MB | Delivered |
| Seedream 5 Litewithdrawn | 1 cr | 42.2s | 2048×2048 · 0.51 MB | Delivered |
| Ideogram V4 Turbowithdrawn | 1 cr | — | — | Provider failure |
| Nano Bananawithdrawn | 5 cr | 16.2s | 1024×1024 · 1.3 MB | Delivered |
| DALL E 3withdrawn | 5 cr | — | — | Provider failure |
| Qwen Image 2withdrawn | 5 cr | 16s | 2048×2048 · 4.6 MB | Delivered |
| GPT Image 1.5withdrawn | 20 cr | 48.4s | 1024×1024 · 1.6 MB | Delivered |
| Midjourney V7withdrawn | 20 cr | 54.5s | 1024×1024 · 1.6 MB | Delivered |
| Seedream 5 Prowithdrawn | 20 cr | 61.3s | 2000×2000 · 4.2 MB | Delivered |
| Flux 2 Maxwithdrawn | 20 cr | 22.6s | 1008×1008 · 0.05 MB | Delivered |
| Ideogram V4 Qualitywithdrawn | 20 cr | — | — | Provider failure |
3 of the 11 engines took the job and then returned a failure from the provider. Credits are refunded automatically when a render fails, so those rows cost nothing — they stay in the table because that is the part of a benchmark nobody else publishes. 11 of them have since been withdrawn from the studio altogether: the failure was the provider's, it kept happening, and no backup renderer carries them — so they are no longer offered rather than left in the picker to fail again.
Every output
The actual files, unretouched.
First and only attempt from each engine. Open any one to see its full run page.
1 cr · 9.5s · 1024×1024 · 0.04 MB
1 cr · 42.2s · 2048×2048 · 0.51 MB
5 cr · 16.2s · 1024×1024 · 1.3 MB
5 cr · 16s · 2048×2048 · 4.6 MB
20 cr · 48.4s · 1024×1024 · 1.6 MB
20 cr · 54.5s · 1024×1024 · 1.6 MB
20 cr · 61.3s · 2000×2000 · 4.2 MB
20 cr · 22.6s · 1008×1008 · 0.05 MB
What it showed
Reading the run.
- The 1-credit Seedream 5 Lite came back at 2048×2048. The 20-credit Flux 2 Max came back at 1008×1008 in a 50 KB file. Twenty times the price bought a quarter of the pixels.
- Cheap is fast. Flux 2 Dev delivered in 9.5 seconds for one credit; nothing in the 20-credit tier came back in under 22 seconds, and the slowest took just over a minute.
- Every engine that delivered produced a usable product still. The separation was in how much licence each took with the brief: the two 1-credit engines stuck closest to it, while three of the four 20-credit engines redesigned something — a carabiner and lid loop nobody asked for, a squat tumbler instead of the bottle, a studio sweep instead of overcast daylight.
- One engine moved the shoot: Qwen Image 2 put the bottle on a shoreline with a horizon behind it. Whether that is a bonus or a re-run depends entirely on your brief, which is the argument for drafting cheap before you commit.
- Three of the eleven took the job and then failed — both Ideogram V4 models and DALL·E 3 — twice, on separate days. All three have since been withdrawn from the studio: the failures were the provider's, they kept repeating, and no backup renderer carries those models, so there was nothing to be gained by leaving them in the picker. That is why those three rows have no model page to open.
- File size follows compression, not quality: two engines returned under 60 KB where others returned over 4 MB for a comparable frame. That matters if you are uploading to a marketplace with a minimum file-size rule.
FAQ
What people ask.
For this brief, price did not predict the result. The two 1-credit engines stuck closest to the product described, one of them returned the largest frame in the whole run at 2048×2048, and three of the four 20-credit engines redesigned something — hardware the brief never mentioned, a different bottle shape, a studio backdrop instead of overcast daylight. That will not hold for every job: typography, faces and character consistency separate these engines far more than a product still does. The point of publishing the run is that you can look at the eight outputs and decide for your own work instead of trusting a ranking.
Between 1 and 20 credits a render in this studio — a fraction of a cent to about six cents at Starter pricing. Eight delivered images across the whole price range came to 92 credits, about 26 cents, which is why drafting on the cheap engines and finishing once on an expensive one is the only sensible way to work.
Three engines accepted the job and then returned a failure from the upstream provider, twice in a row, on different days. Credits are refunded automatically when a render fails, so the run cost nothing for those rows — but they stay in the table, because a model that does not work is the most useful thing a benchmark can tell you.
No. One attempt per model, no re-rolls, no prompt tuning per engine, and every output published to the public wall where anyone can open it. The same prompt and settings are printed above the table so you can run it yourself and compare.
Method
How this was run, so you can argue with it.
Each model received the identical prompt above through the same studio a customer uses, one model at a time so nothing queued behind anything else. Time is wall-clock from the moment the request was accepted to the moment a playable file existed — provider queue included, because that is the wait you actually experience. It is a single run per model, not an average, and a second run would land somewhere slightly different.
Every delivered file is shown unedited and keeps its own run page, which is what the thumbnails above link to. Nothing was upscaled, re-rolled, colour-corrected or re-prompted per engine. The credit figures are this studio's list prices, and the total is what the delivered renders cost — a render that fails refunds itself, so the failures above are free in money and expensive only in time.
Run your own brief instead
Every model still offered in this table is in the same picker, on one balance. Free account, 22 credits a day, no card.