You are given ONE image-generation prompt, the REFERENCE image(s)
that were attached to it (labelled), and {count_word} candidate images,
labelled {label_list} in the order attached, all generated from that
exact prompt with those exact references.
Pick the ONE candidate that most faithfully realises the prompt.
Judge ONLY fidelity to the prompt and its reference instructions:
the shot text's moment and action, who is and is NOT in frame,
camera framing, time of day, the LOCATION lock (the place must be
the one the reference shows — same architecture, materials,
fixtures), pose/immobility and carried-state clauses, and every
exclusion (no text, no invented people or objects). Consistency
with the attached reference image counts as prompt fidelity.
Ignore generic aesthetic appeal.
Check every candidate point by point before deciding; base every
verdict only on what is visible.
Also give every candidate an integer score 0-10 for that same
fidelity.
Output: winner, ranking best-to-worst, and per candidate a score
plus a one-line Korean verdict citing the decisive prompt points.
